Strip a leading byte order mark before parsing - #801

Merged
hsbt merged 4 commits into
masterfrom
claude/priceless-cannon-38818f
Jun 12, 2026
Merged

Strip a leading byte order mark before parsing#801
hsbt merged 4 commits into
masterfrom
claude/priceless-cannon-38818f

Conversation

@hsbt

@hsbthsbt commented Jun 12, 2026

Copy link
Copy Markdown
Member

Fixes#331.

YAML.load("\uFEFFa: b\nc: d") returns {"a" => "b"}, silently dropping everything after the first newline, and Psych.parse_stream raises Psych::SyntaxError on the same input. A BOM at the start of the stream is legal per YAML 1.2 §5.2.

Psych tells libyaml the input encoding whenever it is known, so libyaml's reader-level BOM stripping, which only runs during encoding auto-detection, never happens. The scanner skips the BOM instead but counts it as a first-line column, so every token on the first line shifts one column right and a root-level block mapping terminates at the second line. Reported upstream as yaml/libyaml#334.

This strips the BOM at the Ruby level before the input reaches libyaml: the first character of UTF-8/UTF-16 strings, and of seekable IOs whose external encoding is one of those. Binary strings and IOs are left untouched since they go through libyaml's auto-detection, which already handles the BOM correctly. JRuby is unaffected (snakeyaml-engine excludes the BOM from column counting) and the new tests pass on both implementations.

🤖 Generated with Claude Code

libyaml only discounts the BOM when it detects the stream encoding by
itself. Psych passes the encoding explicitly whenever it is known, and
on that path libyaml counts the BOM as a first-line character, shifting
every token on the first line one column right and silently terminating
a block mapping at the second line.
#331
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
CopilotAI review requested due to automatic review settings June 12, 2026 01:25

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR fixes YAML parsing when an input stream begins with a Unicode byte order mark (BOM), aligning behavior with YAML 1.2 by ensuring BOMs don’t shift token columns and break multi-line root-level mappings.

Changes:

  • Strip a leading BOM in Psych::Parser#parse for UTF-8/UTF-16 inputs before handing data to libyaml.
  • Add regression tests for Psych.load, Psych.parse_stream, and Psych::Parser#parse covering multi-line mappings with BOM (UTF-8, UTF-16, and IO cases).

Reviewed changes

Copilot reviewed 3 out of 3 changed files in this pull request and generated 1 comment.

FileDescription
lib/psych/parser.rbAdds Ruby-level BOM stripping before invoking the native libyaml parser.
test/psych/test_psych.rbAdds high-level regression tests for load and parse_stream with a leading BOM.
test/psych/test_parser.rbAdds parser-level BOM regression tests (UTF-8/UTF-16/IO) and a helper for scalar extraction.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment threadlib/psych/parser.rb
Comment on lines +80 to +83
if String === yaml
bom = BOM[yaml.encoding]
return yaml[1..-1] if bom && yaml.start_with?(bom)
elsif yaml.respond_to?(:read) && yaml.respond_to?(:external_encoding) &&
hsbtand others added 3 commits June 12, 2026 11:01
If pos succeeded but the later seek failed, the rescue silently
discarded the bytes read to check for a BOM. Only the initial pos call
is expected to fail, for non-seekable IOs, before anything is consumed.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The C extension transcodes UTF-32 strings to UTF-8 with the BOM
preserved, so they were truncated the same way as UTF-8 input.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@hsbt
hsbt merged commit a1dcb86 into masterJun 12, 2026
164 checks passed
@hsbt
hsbt deleted the claude/priceless-cannon-38818f branch June 12, 2026 02:53
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

YAML parser stops processing at the first newline when a byte order mark is present

2 participants

@hsbt
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Strip a leading byte order mark before parsing - #801

Merged
hsbt merged 4 commits into
masterfrom
claude/priceless-cannon-38818f
Jun 12, 2026
Merged

Strip a leading byte order mark before parsing#801
hsbt merged 4 commits into
masterfrom
claude/priceless-cannon-38818f

Conversation

@hsbt

@hsbthsbt commented Jun 12, 2026

Copy link
Copy Markdown
Member

Fixes#331.

YAML.load("\uFEFFa: b\nc: d") returns {"a" => "b"}, silently dropping everything after the first newline, and Psych.parse_stream raises Psych::SyntaxError on the same input. A BOM at the start of the stream is legal per YAML 1.2 §5.2.

Psych tells libyaml the input encoding whenever it is known, so libyaml's reader-level BOM stripping, which only runs during encoding auto-detection, never happens. The scanner skips the BOM instead but counts it as a first-line column, so every token on the first line shifts one column right and a root-level block mapping terminates at the second line. Reported upstream as yaml/libyaml#334.

This strips the BOM at the Ruby level before the input reaches libyaml: the first character of UTF-8/UTF-16 strings, and of seekable IOs whose external encoding is one of those. Binary strings and IOs are left untouched since they go through libyaml's auto-detection, which already handles the BOM correctly. JRuby is unaffected (snakeyaml-engine excludes the BOM from column counting) and the new tests pass on both implementations.

🤖 Generated with Claude Code

libyaml only discounts the BOM when it detects the stream encoding by
itself. Psych passes the encoding explicitly whenever it is known, and
on that path libyaml counts the BOM as a first-line character, shifting
every token on the first line one column right and silently terminating
a block mapping at the second line.
#331
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
CopilotAI review requested due to automatic review settings June 12, 2026 01:25

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR fixes YAML parsing when an input stream begins with a Unicode byte order mark (BOM), aligning behavior with YAML 1.2 by ensuring BOMs don’t shift token columns and break multi-line root-level mappings.

Changes:

  • Strip a leading BOM in Psych::Parser#parse for UTF-8/UTF-16 inputs before handing data to libyaml.
  • Add regression tests for Psych.load, Psych.parse_stream, and Psych::Parser#parse covering multi-line mappings with BOM (UTF-8, UTF-16, and IO cases).

Reviewed changes

Copilot reviewed 3 out of 3 changed files in this pull request and generated 1 comment.

FileDescription
lib/psych/parser.rbAdds Ruby-level BOM stripping before invoking the native libyaml parser.
test/psych/test_psych.rbAdds high-level regression tests for load and parse_stream with a leading BOM.
test/psych/test_parser.rbAdds parser-level BOM regression tests (UTF-8/UTF-16/IO) and a helper for scalar extraction.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment threadlib/psych/parser.rb
Comment on lines +80 to +83
if String === yaml
bom = BOM[yaml.encoding]
return yaml[1..-1] if bom && yaml.start_with?(bom)
elsif yaml.respond_to?(:read) && yaml.respond_to?(:external_encoding) &&
hsbtand others added 3 commits June 12, 2026 11:01
If pos succeeded but the later seek failed, the rescue silently
discarded the bytes read to check for a BOM. Only the initial pos call
is expected to fail, for non-seekable IOs, before anything is consumed.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The C extension transcodes UTF-32 strings to UTF-8 with the BOM
preserved, so they were truncated the same way as UTF-8 input.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@hsbt
hsbt merged commit a1dcb86 into masterJun 12, 2026
164 checks passed
@hsbt
hsbt deleted the claude/priceless-cannon-38818f branch June 12, 2026 02:53
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

YAML parser stops processing at the first newline when a byte order mark is present

2 participants

@hsbt
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Strip a leading byte order mark before parsing - #801

Merged
hsbt merged 4 commits into
masterfrom
claude/priceless-cannon-38818f
Jun 12, 2026
Merged

Strip a leading byte order mark before parsing#801
hsbt merged 4 commits into
masterfrom
claude/priceless-cannon-38818f

Conversation

@hsbt

@hsbthsbt commented Jun 12, 2026

Copy link
Copy Markdown
Member

Fixes#331.

YAML.load("\uFEFFa: b\nc: d") returns {"a" => "b"}, silently dropping everything after the first newline, and Psych.parse_stream raises Psych::SyntaxError on the same input. A BOM at the start of the stream is legal per YAML 1.2 §5.2.

Psych tells libyaml the input encoding whenever it is known, so libyaml's reader-level BOM stripping, which only runs during encoding auto-detection, never happens. The scanner skips the BOM instead but counts it as a first-line column, so every token on the first line shifts one column right and a root-level block mapping terminates at the second line. Reported upstream as yaml/libyaml#334.

This strips the BOM at the Ruby level before the input reaches libyaml: the first character of UTF-8/UTF-16 strings, and of seekable IOs whose external encoding is one of those. Binary strings and IOs are left untouched since they go through libyaml's auto-detection, which already handles the BOM correctly. JRuby is unaffected (snakeyaml-engine excludes the BOM from column counting) and the new tests pass on both implementations.

🤖 Generated with Claude Code

libyaml only discounts the BOM when it detects the stream encoding by
itself. Psych passes the encoding explicitly whenever it is known, and
on that path libyaml counts the BOM as a first-line character, shifting
every token on the first line one column right and silently terminating
a block mapping at the second line.
#331
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
CopilotAI review requested due to automatic review settings June 12, 2026 01:25

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR fixes YAML parsing when an input stream begins with a Unicode byte order mark (BOM), aligning behavior with YAML 1.2 by ensuring BOMs don’t shift token columns and break multi-line root-level mappings.

Changes:

  • Strip a leading BOM in Psych::Parser#parse for UTF-8/UTF-16 inputs before handing data to libyaml.
  • Add regression tests for Psych.load, Psych.parse_stream, and Psych::Parser#parse covering multi-line mappings with BOM (UTF-8, UTF-16, and IO cases).

Reviewed changes

Copilot reviewed 3 out of 3 changed files in this pull request and generated 1 comment.

FileDescription
lib/psych/parser.rbAdds Ruby-level BOM stripping before invoking the native libyaml parser.
test/psych/test_psych.rbAdds high-level regression tests for load and parse_stream with a leading BOM.
test/psych/test_parser.rbAdds parser-level BOM regression tests (UTF-8/UTF-16/IO) and a helper for scalar extraction.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment threadlib/psych/parser.rb
Comment on lines +80 to +83
if String === yaml
bom = BOM[yaml.encoding]
return yaml[1..-1] if bom && yaml.start_with?(bom)
elsif yaml.respond_to?(:read) && yaml.respond_to?(:external_encoding) &&
hsbtand others added 3 commits June 12, 2026 11:01
If pos succeeded but the later seek failed, the rescue silently
discarded the bytes read to check for a BOM. Only the initial pos call
is expected to fail, for non-seekable IOs, before anything is consumed.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The C extension transcodes UTF-32 strings to UTF-8 with the BOM
preserved, so they were truncated the same way as UTF-8 input.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@hsbt
hsbt merged commit a1dcb86 into masterJun 12, 2026
164 checks passed
@hsbt
hsbt deleted the claude/priceless-cannon-38818f branch June 12, 2026 02:53
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

YAML parser stops processing at the first newline when a byte order mark is present

2 participants

@hsbt
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Strip a leading byte order mark before parsing - #801

Merged
hsbt merged 4 commits into
masterfrom
claude/priceless-cannon-38818f
Jun 12, 2026
Merged

Strip a leading byte order mark before parsing#801
hsbt merged 4 commits into
masterfrom
claude/priceless-cannon-38818f

Conversation

@hsbt

@hsbthsbt commented Jun 12, 2026

Copy link
Copy Markdown
Member

Fixes#331.

YAML.load("\uFEFFa: b\nc: d") returns {"a" => "b"}, silently dropping everything after the first newline, and Psych.parse_stream raises Psych::SyntaxError on the same input. A BOM at the start of the stream is legal per YAML 1.2 §5.2.

Psych tells libyaml the input encoding whenever it is known, so libyaml's reader-level BOM stripping, which only runs during encoding auto-detection, never happens. The scanner skips the BOM instead but counts it as a first-line column, so every token on the first line shifts one column right and a root-level block mapping terminates at the second line. Reported upstream as yaml/libyaml#334.

This strips the BOM at the Ruby level before the input reaches libyaml: the first character of UTF-8/UTF-16 strings, and of seekable IOs whose external encoding is one of those. Binary strings and IOs are left untouched since they go through libyaml's auto-detection, which already handles the BOM correctly. JRuby is unaffected (snakeyaml-engine excludes the BOM from column counting) and the new tests pass on both implementations.

🤖 Generated with Claude Code

libyaml only discounts the BOM when it detects the stream encoding by
itself. Psych passes the encoding explicitly whenever it is known, and
on that path libyaml counts the BOM as a first-line character, shifting
every token on the first line one column right and silently terminating
a block mapping at the second line.
#331
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
CopilotAI review requested due to automatic review settings June 12, 2026 01:25

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR fixes YAML parsing when an input stream begins with a Unicode byte order mark (BOM), aligning behavior with YAML 1.2 by ensuring BOMs don’t shift token columns and break multi-line root-level mappings.

Changes:

  • Strip a leading BOM in Psych::Parser#parse for UTF-8/UTF-16 inputs before handing data to libyaml.
  • Add regression tests for Psych.load, Psych.parse_stream, and Psych::Parser#parse covering multi-line mappings with BOM (UTF-8, UTF-16, and IO cases).

Reviewed changes

Copilot reviewed 3 out of 3 changed files in this pull request and generated 1 comment.

FileDescription
lib/psych/parser.rbAdds Ruby-level BOM stripping before invoking the native libyaml parser.
test/psych/test_psych.rbAdds high-level regression tests for load and parse_stream with a leading BOM.
test/psych/test_parser.rbAdds parser-level BOM regression tests (UTF-8/UTF-16/IO) and a helper for scalar extraction.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment threadlib/psych/parser.rb
Comment on lines +80 to +83
if String === yaml
bom = BOM[yaml.encoding]
return yaml[1..-1] if bom && yaml.start_with?(bom)
elsif yaml.respond_to?(:read) && yaml.respond_to?(:external_encoding) &&
hsbtand others added 3 commits June 12, 2026 11:01
If pos succeeded but the later seek failed, the rescue silently
discarded the bytes read to check for a BOM. Only the initial pos call
is expected to fail, for non-seekable IOs, before anything is consumed.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The C extension transcodes UTF-32 strings to UTF-8 with the BOM
preserved, so they were truncated the same way as UTF-8 input.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@hsbt
hsbt merged commit a1dcb86 into masterJun 12, 2026
164 checks passed
@hsbt
hsbt deleted the claude/priceless-cannon-38818f branch June 12, 2026 02:53
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

YAML parser stops processing at the first newline when a byte order mark is present

2 participants

@hsbt
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Strip a leading byte order mark before parsing - #801

Merged
hsbt merged 4 commits into
masterfrom
claude/priceless-cannon-38818f
Jun 12, 2026
Merged

Strip a leading byte order mark before parsing#801
hsbt merged 4 commits into
masterfrom
claude/priceless-cannon-38818f

Conversation

@hsbt

@hsbthsbt commented Jun 12, 2026

Copy link
Copy Markdown
Member

Fixes#331.

YAML.load("\uFEFFa: b\nc: d") returns {"a" => "b"}, silently dropping everything after the first newline, and Psych.parse_stream raises Psych::SyntaxError on the same input. A BOM at the start of the stream is legal per YAML 1.2 §5.2.

Psych tells libyaml the input encoding whenever it is known, so libyaml's reader-level BOM stripping, which only runs during encoding auto-detection, never happens. The scanner skips the BOM instead but counts it as a first-line column, so every token on the first line shifts one column right and a root-level block mapping terminates at the second line. Reported upstream as yaml/libyaml#334.

This strips the BOM at the Ruby level before the input reaches libyaml: the first character of UTF-8/UTF-16 strings, and of seekable IOs whose external encoding is one of those. Binary strings and IOs are left untouched since they go through libyaml's auto-detection, which already handles the BOM correctly. JRuby is unaffected (snakeyaml-engine excludes the BOM from column counting) and the new tests pass on both implementations.

🤖 Generated with Claude Code

libyaml only discounts the BOM when it detects the stream encoding by
itself. Psych passes the encoding explicitly whenever it is known, and
on that path libyaml counts the BOM as a first-line character, shifting
every token on the first line one column right and silently terminating
a block mapping at the second line.
#331
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
CopilotAI review requested due to automatic review settings June 12, 2026 01:25

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR fixes YAML parsing when an input stream begins with a Unicode byte order mark (BOM), aligning behavior with YAML 1.2 by ensuring BOMs don’t shift token columns and break multi-line root-level mappings.

Changes:

  • Strip a leading BOM in Psych::Parser#parse for UTF-8/UTF-16 inputs before handing data to libyaml.
  • Add regression tests for Psych.load, Psych.parse_stream, and Psych::Parser#parse covering multi-line mappings with BOM (UTF-8, UTF-16, and IO cases).

Reviewed changes

Copilot reviewed 3 out of 3 changed files in this pull request and generated 1 comment.

FileDescription
lib/psych/parser.rbAdds Ruby-level BOM stripping before invoking the native libyaml parser.
test/psych/test_psych.rbAdds high-level regression tests for load and parse_stream with a leading BOM.
test/psych/test_parser.rbAdds parser-level BOM regression tests (UTF-8/UTF-16/IO) and a helper for scalar extraction.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment threadlib/psych/parser.rb
Comment on lines +80 to +83
if String === yaml
bom = BOM[yaml.encoding]
return yaml[1..-1] if bom && yaml.start_with?(bom)
elsif yaml.respond_to?(:read) && yaml.respond_to?(:external_encoding) &&
hsbtand others added 3 commits June 12, 2026 11:01
If pos succeeded but the later seek failed, the rescue silently
discarded the bytes read to check for a BOM. Only the initial pos call
is expected to fail, for non-seekable IOs, before anything is consumed.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The C extension transcodes UTF-32 strings to UTF-8 with the BOM
preserved, so they were truncated the same way as UTF-8 input.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@hsbt
hsbt merged commit a1dcb86 into masterJun 12, 2026
164 checks passed
@hsbt
hsbt deleted the claude/priceless-cannon-38818f branch June 12, 2026 02:53
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

YAML parser stops processing at the first newline when a byte order mark is present

2 participants

@hsbt
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Strip a leading byte order mark before parsing - #801

Merged
hsbt merged 4 commits into
masterfrom
claude/priceless-cannon-38818f
Jun 12, 2026
Merged

Strip a leading byte order mark before parsing#801
hsbt merged 4 commits into
masterfrom
claude/priceless-cannon-38818f

Conversation

@hsbt

@hsbthsbt commented Jun 12, 2026

Copy link
Copy Markdown
Member

Fixes#331.

YAML.load("\uFEFFa: b\nc: d") returns {"a" => "b"}, silently dropping everything after the first newline, and Psych.parse_stream raises Psych::SyntaxError on the same input. A BOM at the start of the stream is legal per YAML 1.2 §5.2.

Psych tells libyaml the input encoding whenever it is known, so libyaml's reader-level BOM stripping, which only runs during encoding auto-detection, never happens. The scanner skips the BOM instead but counts it as a first-line column, so every token on the first line shifts one column right and a root-level block mapping terminates at the second line. Reported upstream as yaml/libyaml#334.

This strips the BOM at the Ruby level before the input reaches libyaml: the first character of UTF-8/UTF-16 strings, and of seekable IOs whose external encoding is one of those. Binary strings and IOs are left untouched since they go through libyaml's auto-detection, which already handles the BOM correctly. JRuby is unaffected (snakeyaml-engine excludes the BOM from column counting) and the new tests pass on both implementations.

🤖 Generated with Claude Code

libyaml only discounts the BOM when it detects the stream encoding by
itself. Psych passes the encoding explicitly whenever it is known, and
on that path libyaml counts the BOM as a first-line character, shifting
every token on the first line one column right and silently terminating
a block mapping at the second line.
#331
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
CopilotAI review requested due to automatic review settings June 12, 2026 01:25

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR fixes YAML parsing when an input stream begins with a Unicode byte order mark (BOM), aligning behavior with YAML 1.2 by ensuring BOMs don’t shift token columns and break multi-line root-level mappings.

Changes:

  • Strip a leading BOM in Psych::Parser#parse for UTF-8/UTF-16 inputs before handing data to libyaml.
  • Add regression tests for Psych.load, Psych.parse_stream, and Psych::Parser#parse covering multi-line mappings with BOM (UTF-8, UTF-16, and IO cases).

Reviewed changes

Copilot reviewed 3 out of 3 changed files in this pull request and generated 1 comment.

FileDescription
lib/psych/parser.rbAdds Ruby-level BOM stripping before invoking the native libyaml parser.
test/psych/test_psych.rbAdds high-level regression tests for load and parse_stream with a leading BOM.
test/psych/test_parser.rbAdds parser-level BOM regression tests (UTF-8/UTF-16/IO) and a helper for scalar extraction.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment threadlib/psych/parser.rb
Comment on lines +80 to +83
if String === yaml
bom = BOM[yaml.encoding]
return yaml[1..-1] if bom && yaml.start_with?(bom)
elsif yaml.respond_to?(:read) && yaml.respond_to?(:external_encoding) &&
hsbtand others added 3 commits June 12, 2026 11:01
If pos succeeded but the later seek failed, the rescue silently
discarded the bytes read to check for a BOM. Only the initial pos call
is expected to fail, for non-seekable IOs, before anything is consumed.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The C extension transcodes UTF-32 strings to UTF-8 with the BOM
preserved, so they were truncated the same way as UTF-8 input.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@hsbt
hsbt merged commit a1dcb86 into masterJun 12, 2026
164 checks passed
@hsbt
hsbt deleted the claude/priceless-cannon-38818f branch June 12, 2026 02:53
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

YAML parser stops processing at the first newline when a byte order mark is present

2 participants

@hsbt
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Strip a leading byte order mark before parsing - #801

Merged
hsbt merged 4 commits into
masterfrom
claude/priceless-cannon-38818f
Jun 12, 2026
Merged

Strip a leading byte order mark before parsing#801
hsbt merged 4 commits into
masterfrom
claude/priceless-cannon-38818f

Conversation

@hsbt

@hsbthsbt commented Jun 12, 2026

Copy link
Copy Markdown
Member

Fixes#331.

YAML.load("\uFEFFa: b\nc: d") returns {"a" => "b"}, silently dropping everything after the first newline, and Psych.parse_stream raises Psych::SyntaxError on the same input. A BOM at the start of the stream is legal per YAML 1.2 §5.2.

Psych tells libyaml the input encoding whenever it is known, so libyaml's reader-level BOM stripping, which only runs during encoding auto-detection, never happens. The scanner skips the BOM instead but counts it as a first-line column, so every token on the first line shifts one column right and a root-level block mapping terminates at the second line. Reported upstream as yaml/libyaml#334.

This strips the BOM at the Ruby level before the input reaches libyaml: the first character of UTF-8/UTF-16 strings, and of seekable IOs whose external encoding is one of those. Binary strings and IOs are left untouched since they go through libyaml's auto-detection, which already handles the BOM correctly. JRuby is unaffected (snakeyaml-engine excludes the BOM from column counting) and the new tests pass on both implementations.

🤖 Generated with Claude Code

libyaml only discounts the BOM when it detects the stream encoding by
itself. Psych passes the encoding explicitly whenever it is known, and
on that path libyaml counts the BOM as a first-line character, shifting
every token on the first line one column right and silently terminating
a block mapping at the second line.
#331
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
CopilotAI review requested due to automatic review settings June 12, 2026 01:25

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR fixes YAML parsing when an input stream begins with a Unicode byte order mark (BOM), aligning behavior with YAML 1.2 by ensuring BOMs don’t shift token columns and break multi-line root-level mappings.

Changes:

  • Strip a leading BOM in Psych::Parser#parse for UTF-8/UTF-16 inputs before handing data to libyaml.
  • Add regression tests for Psych.load, Psych.parse_stream, and Psych::Parser#parse covering multi-line mappings with BOM (UTF-8, UTF-16, and IO cases).

Reviewed changes

Copilot reviewed 3 out of 3 changed files in this pull request and generated 1 comment.

FileDescription
lib/psych/parser.rbAdds Ruby-level BOM stripping before invoking the native libyaml parser.
test/psych/test_psych.rbAdds high-level regression tests for load and parse_stream with a leading BOM.
test/psych/test_parser.rbAdds parser-level BOM regression tests (UTF-8/UTF-16/IO) and a helper for scalar extraction.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment threadlib/psych/parser.rb
Comment on lines +80 to +83
if String === yaml
bom = BOM[yaml.encoding]
return yaml[1..-1] if bom && yaml.start_with?(bom)
elsif yaml.respond_to?(:read) && yaml.respond_to?(:external_encoding) &&
hsbtand others added 3 commits June 12, 2026 11:01
If pos succeeded but the later seek failed, the rescue silently
discarded the bytes read to check for a BOM. Only the initial pos call
is expected to fail, for non-seekable IOs, before anything is consumed.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The C extension transcodes UTF-32 strings to UTF-8 with the BOM
preserved, so they were truncated the same way as UTF-8 input.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@hsbt
hsbt merged commit a1dcb86 into masterJun 12, 2026
164 checks passed
@hsbt
hsbt deleted the claude/priceless-cannon-38818f branch June 12, 2026 02:53
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

YAML parser stops processing at the first newline when a byte order mark is present

2 participants

@hsbt
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Strip a leading byte order mark before parsing - #801

Merged
hsbt merged 4 commits into
masterfrom
claude/priceless-cannon-38818f
Jun 12, 2026
Merged

Strip a leading byte order mark before parsing#801
hsbt merged 4 commits into
masterfrom
claude/priceless-cannon-38818f

Conversation

@hsbt

@hsbthsbt commented Jun 12, 2026

Copy link
Copy Markdown
Member

Fixes#331.

YAML.load("\uFEFFa: b\nc: d") returns {"a" => "b"}, silently dropping everything after the first newline, and Psych.parse_stream raises Psych::SyntaxError on the same input. A BOM at the start of the stream is legal per YAML 1.2 §5.2.

Psych tells libyaml the input encoding whenever it is known, so libyaml's reader-level BOM stripping, which only runs during encoding auto-detection, never happens. The scanner skips the BOM instead but counts it as a first-line column, so every token on the first line shifts one column right and a root-level block mapping terminates at the second line. Reported upstream as yaml/libyaml#334.

This strips the BOM at the Ruby level before the input reaches libyaml: the first character of UTF-8/UTF-16 strings, and of seekable IOs whose external encoding is one of those. Binary strings and IOs are left untouched since they go through libyaml's auto-detection, which already handles the BOM correctly. JRuby is unaffected (snakeyaml-engine excludes the BOM from column counting) and the new tests pass on both implementations.

🤖 Generated with Claude Code

libyaml only discounts the BOM when it detects the stream encoding by
itself. Psych passes the encoding explicitly whenever it is known, and
on that path libyaml counts the BOM as a first-line character, shifting
every token on the first line one column right and silently terminating
a block mapping at the second line.
#331
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
CopilotAI review requested due to automatic review settings June 12, 2026 01:25

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR fixes YAML parsing when an input stream begins with a Unicode byte order mark (BOM), aligning behavior with YAML 1.2 by ensuring BOMs don’t shift token columns and break multi-line root-level mappings.

Changes:

  • Strip a leading BOM in Psych::Parser#parse for UTF-8/UTF-16 inputs before handing data to libyaml.
  • Add regression tests for Psych.load, Psych.parse_stream, and Psych::Parser#parse covering multi-line mappings with BOM (UTF-8, UTF-16, and IO cases).

Reviewed changes

Copilot reviewed 3 out of 3 changed files in this pull request and generated 1 comment.

FileDescription
lib/psych/parser.rbAdds Ruby-level BOM stripping before invoking the native libyaml parser.
test/psych/test_psych.rbAdds high-level regression tests for load and parse_stream with a leading BOM.
test/psych/test_parser.rbAdds parser-level BOM regression tests (UTF-8/UTF-16/IO) and a helper for scalar extraction.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment threadlib/psych/parser.rb
Comment on lines +80 to +83
if String === yaml
bom = BOM[yaml.encoding]
return yaml[1..-1] if bom && yaml.start_with?(bom)
elsif yaml.respond_to?(:read) && yaml.respond_to?(:external_encoding) &&
hsbtand others added 3 commits June 12, 2026 11:01
If pos succeeded but the later seek failed, the rescue silently
discarded the bytes read to check for a BOM. Only the initial pos call
is expected to fail, for non-seekable IOs, before anything is consumed.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The C extension transcodes UTF-32 strings to UTF-8 with the BOM
preserved, so they were truncated the same way as UTF-8 input.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@hsbt
hsbt merged commit a1dcb86 into masterJun 12, 2026
164 checks passed
@hsbt
hsbt deleted the claude/priceless-cannon-38818f branch June 12, 2026 02:53
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

YAML parser stops processing at the first newline when a byte order mark is present

2 participants

@hsbt