[BugFix] Fix unexpected shift of extraction for rex with nested capture groups in named groups - #4641

Merged
RyanL1997 merged 7 commits into
opensearch-project:mainfrom
RyanL1997:rex-extract-fix
Oct 29, 2025
Merged

[BugFix] Fix unexpected shift of extraction for rex with nested capture groups in named groups #4641
RyanL1997 merged 7 commits into
opensearch-project:mainfrom
RyanL1997:rex-extract-fix

Conversation

@RyanL1997

@RyanL1997RyanL1997 commented Oct 22, 2025

Copy link
Copy Markdown
Collaborator

Description

Fix unexpected shift of extraction for rex with nested capture groups in named groups

The rex command in PPL had a critical bug when using named capture groups that contained nested unnamed groups. This caused extracted field values to shift by one position, producing incorrect results.

  • Root Cause: Code used sequential indices 1, 2, 3... but nested groups create non-sequential indices 1, 3, 5...
  • Solution: Bypass index calculation entirely by using Java's native named group extraction (matcher.group(groupName))

Example of the Bug

Query:

curl -X POST "localhost:9200/_plugins/_ppl" \
-H "Content-Type: application/json" \
-d '{ "query": "source=accounts | rex field=email \"(?<user>(amber|hattie|nanette)[a-z]*)@(?<domain>(pyrami|netagy|quility))\\.(?<tld>(com|org))\" | fields user, domain, tld | head 1" }'

Expected Result (correct):

["amberduke", "pyrami", "com"]

Actual Result (wrong):

["amberduke", "amber", "pyrami"]

Root Cause

When Java's regex engine processes the pattern (?<user>(amber|hattie))[a-z]*, it assigns group numbers to ALL capture groups (named and unnamed):

Pattern: (?<user>(amber|hattie))[a-z]*@(?<domain>(pyrami|netagy))\.(?<tld>(com|org))

Group Assignment:

  • Group 0: Entire match
  • Group 1: (?<user>...) ← Named group "user"
  • Group 2: (amber|hattie) ← Unnamed nested group
  • Group 3: (?<domain>...) ← Named group "domain"
  • Group 4: (pyrami|netagy) ← Unnamed nested group
  • Group 5: (?<tld>...) ← Named group "tld"
  • Group 6: (com|org) ← Unnamed nested group

Before the fix, the bug is in CalciteRelNodeVisitor.java (lines 265-321). The code does:

List<String> namedGroups = RegexCommonUtils.getNamedGroupCandidates(patternStr);
// namedGroups = ["user", "domain", "tld"]for (inti = 0; i < namedGroups.size(); i++) {
extractCall = PPLFuncImpTable.INSTANCE.resolve(
context.rexBuilder,
BuiltinFunctionName.REX_EXTRACT,
fieldRex,
context.rexBuilder.makeLiteral(patternStr),
context.relBuilder.literal(i + 1)); // ← WRONG: Assumes sequential named groups// ...
}

The code assumes named groups are at indices 1, 2, 3, ... but the actual indices are 1, 3, 5, ... due to the unnamed nested groups.

With the above buggy logic:

  • REX_EXTRACT(field, pattern, 1) → Gets Group 1 (?<user>...) = "amberduke" → CORRECT
  • REX_EXTRACT(field, pattern, 2) → Gets Group 2 (amber|hattie) = "amber" → WRONG
  • REX_EXTRACT(field, pattern, 3) → Gets Group 3 (?<domain>...) = "pyrami" → WRONG

The second and third extractions are off by one group because they hit the unnamed nested groups.

LogicalProject(
user=[REX_EXTRACT($7, '(?<user>(amber|hattie))[a-z]*@(?<domain>(pyrami|netagy))\.(?<tld>(com|org))', 1)],
domain=[REX_EXTRACT($7, '(?<user>(amber|hattie))[a-z]*@(?<domain>(pyrami|netagy))\.(?<tld>(com|org))', 2)], -- Wrong!
tld=[REX_EXTRACT($7, '(?<user>(amber|hattie))[a-z]*@(?<domain>(pyrami|netagy))\.(?<tld>(com|org))', 3)] -- Wrong!
)

Related Issues

Check List

  • New functionality includes testing.
  • New functionality has been documented.
  • New functionality has javadoc added.
  • New functionality has a user manual doc added.
  • New PPL command checklist all confirmed.
  • API changes companion pull request created.
  • Commits are signed per the DCO using --signoff or -s.
  • Public documentation issue/PR created.

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.
For more information on following Developer Certificate of Origin and signing off your commits, please check here.

matchCount++;
} else {
// If extractor returns null, it might indicate an error (like invalid group name)
// Stop processing to avoid infinite loop

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm currently thinking about adding an error handling here

Signed-off-by: Jialiang Liang <jiallian@amazon.com>
Signed-off-by: Jialiang Liang <jiallian@amazon.com>
Signed-off-by: Jialiang Liang <jiallian@amazon.com>
Signed-off-by: Jialiang Liang <jiallian@amazon.com>
Signed-off-by: Jialiang Liang <jiallian@amazon.com>
Signed-off-by: Jialiang Liang <jiallian@amazon.com>
@RyanL1997RyanL1997 changed the title [BugFix] Fix the off-by-one error for rex with nested capture groups in named groups [BugFix] Fix unexpected shift of extraction for rex with nested capture groups in named groups Oct 25, 2025
Swiddis
Swiddis previously approved these changes Oct 28, 2025
fieldRex,
context.rexBuilder.makeLiteral(patternStr),
context.relBuilder.literal(i + 1),
context.rexBuilder.makeLiteral(groupName),

@dai-chendai-chenOct 28, 2025

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I found there is a namedGroups() API in Matcher (since JDK 20?). If we can get correct index here, we don't need to modify the UDFs below? Alternatively we can move capture name -> index logic here from UDFs?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That is correct - the core issue is matching named groups to their correct indices, and Pattern.namedGroups() would be the perfect solution. However, I discovered that we're blocked by a
compatibility constraint:

  • Pattern.namedGroups() was introduced in JDK 20
  • We need backward compatibility with JDK 11/17 - for 2.19-dev

I agree that directly leveraging the Pattern.namedGroups() is the right architectural approach - we should definitely migrate to it when we fully upgrade to JDK 20+. At that point, it would be a simple one-line change in CalciteRelNodeVisitor.

Signed-off-by: Jialiang Liang <jiallian@amazon.com>
"Rex pattern must contain at least one named capture group");
}

// TODO: Once JDK 20+ is supported, consider using Pattern.namedGroups() API for more efficient

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

as for now, I added a TODO here @dai-chen

@RyanL1997
RyanL1997 merged commit 0c1ec27 into opensearch-project:mainOct 29, 2025
51 of 52 checks passed
@RyanL1997
RyanL1997 deleted the rex-extract-fix branch October 29, 2025 00:34
opensearch-trigger-botBot pushed a commit that referenced this pull request Oct 29, 2025
…ture groups in named groups (#4641)
(cherry picked from commit 0c1ec27)
Signed-off-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
RyanL1997 pushed a commit that referenced this pull request Oct 29, 2025
…ture groups in named groups (#4641) (#4692)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
expani pushed a commit to vinaykpud/sql that referenced this pull request Nov 4, 2025
sandeshkr419 added a commit to sandeshkr419/sql that referenced this pull request Dec 3, 2025
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Simeon Widdis <sawiddis@amazon.com>
Co-authored-by: Manasvini B S <manasvis@amazon.com>
Co-authored-by: opensearch-ci-bot <opensearch-infra@amazon.com>
Co-authored-by: Louis Chu <clingzhi@amazon.com>
Co-authored-by: Chen Dai <daichen@amazon.com>
Co-authored-by: Mebsina <cnoramut@gmail.com>
Co-authored-by: Yuanchun Shen <yuanchu@amazon.com>
Co-authored-by: opensearch-trigger-bot[bot] <98922864+opensearch-trigger-bot[bot]@users.noreply.github.com>
Co-authored-by: Kai Huang <105710027+ahkcs@users.noreply.github.com>
Co-authored-by: Peng Huo <penghuo@gmail.com>
Co-authored-by: Alexey Temnikov <alexey.temnikov@improving.com>
Co-authored-by: Riley Jerger <214163063+RileyJergerAmazon@users.noreply.github.com>
Co-authored-by: Tomoyuki MORITA <moritato@amazon.com>
Co-authored-by: Lantao Jin <ltjin@amazon.com>
Co-authored-by: Songkan Tang <songkant@amazon.com>
Co-authored-by: qianheng <qianheng@amazon.com>
Co-authored-by: Simeon Widdis <sawiddis@gmail.com>
Co-authored-by: Xinyuan Lu <xinyual@amazon.com>
Co-authored-by: Jialiang Liang <jiallian@amazon.com>
Co-authored-by: Peter Zhu <zhujiaxi@amazon.com>
Co-authored-by: Vinay Krishna Pudyodu <vinkrish.neo@gmail.com>
Co-authored-by: expani <anijainc@amazon.com>
Co-authored-by: expani1729 <110471048+expani@users.noreply.github.com>
Co-authored-by: Vamsi Manohar <reddyvam@amazon.com>
Co-authored-by: ritvibhatt <53196324+ritvibhatt@users.noreply.github.com>
Co-authored-by: Xinyu Hao <75524174+ishaoxy@users.noreply.github.com>
Co-authored-by: Marc Handalian <marc.handalian@gmail.com>
Co-authored-by: Marc Handalian <handalm@amazon.com>
Fix join type ambiguous issue when specify the join type with sql-like join criteria (opensearch-project#4474)
Fix issue 4441 (opensearch-project#4449)
Fix missing keywordsCanBeId (opensearch-project#4491)
Fix the bug of explicit makeNullLiteral for UDT fields (opensearch-project#4475)
Fix mapping after aggregation push down (opensearch-project#4500)
Fix percentile bug (opensearch-project#4539)
Fix JsonExtractAllFunctionIT failure (opensearch-project#4556)
Fix sort push down into agg after project already pushed (opensearch-project#4546)
Fix push down failure for min/max on derived field (opensearch-project#4572)
Fix compile issue in main (opensearch-project#4608)
Fix filter parsing failure on date fields with non-default format (opensearch-project#4616)
Fix bin nested fields issue (opensearch-project#4606)
Fix: Support Alias Fields in MIN, MAX, FIRST, LAST, and TAKE Aggregations (opensearch-project#4621)
fix rename issue (opensearch-project#4670)
Fixes for `Multisearch` and `Append` command (opensearch-project#4512)
Fix asc/desc keyword behavior for sort command (opensearch-project#4651)
Fix] Fix unexpected shift of extraction for `rex` with nested capture groups in named groups (opensearch-project#4641)
Fix CVE-2025-48924 (opensearch-project#4665)
Fix sub-fields accessing of generated structs (opensearch-project#4683)
Fix] Incorrect Field Index Mapping in AVG to SUM/COUNT Conversion (opensearch-project#15)
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

backport 2.19-devbugSomething isn't workingbugFixPPLPiped processing language

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[BUG] rex command off-by-one error with nested capture groups in named groups

3 participants

@RyanL1997@Swiddis@dai-chen
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

[BugFix] Fix unexpected shift of extraction for rex with nested capture groups in named groups - #4641

Merged
RyanL1997 merged 7 commits into
opensearch-project:mainfrom
RyanL1997:rex-extract-fix
Oct 29, 2025
Merged

[BugFix] Fix unexpected shift of extraction for rex with nested capture groups in named groups #4641
RyanL1997 merged 7 commits into
opensearch-project:mainfrom
RyanL1997:rex-extract-fix

Conversation

@RyanL1997

@RyanL1997RyanL1997 commented Oct 22, 2025

Copy link
Copy Markdown
Collaborator

Description

Fix unexpected shift of extraction for rex with nested capture groups in named groups

The rex command in PPL had a critical bug when using named capture groups that contained nested unnamed groups. This caused extracted field values to shift by one position, producing incorrect results.

  • Root Cause: Code used sequential indices 1, 2, 3... but nested groups create non-sequential indices 1, 3, 5...
  • Solution: Bypass index calculation entirely by using Java's native named group extraction (matcher.group(groupName))

Example of the Bug

Query:

curl -X POST "localhost:9200/_plugins/_ppl" \
-H "Content-Type: application/json" \
-d '{ "query": "source=accounts | rex field=email \"(?<user>(amber|hattie|nanette)[a-z]*)@(?<domain>(pyrami|netagy|quility))\\.(?<tld>(com|org))\" | fields user, domain, tld | head 1" }'

Expected Result (correct):

["amberduke", "pyrami", "com"]

Actual Result (wrong):

["amberduke", "amber", "pyrami"]

Root Cause

When Java's regex engine processes the pattern (?<user>(amber|hattie))[a-z]*, it assigns group numbers to ALL capture groups (named and unnamed):

Pattern: (?<user>(amber|hattie))[a-z]*@(?<domain>(pyrami|netagy))\.(?<tld>(com|org))

Group Assignment:

  • Group 0: Entire match
  • Group 1: (?<user>...) ← Named group "user"
  • Group 2: (amber|hattie) ← Unnamed nested group
  • Group 3: (?<domain>...) ← Named group "domain"
  • Group 4: (pyrami|netagy) ← Unnamed nested group
  • Group 5: (?<tld>...) ← Named group "tld"
  • Group 6: (com|org) ← Unnamed nested group

Before the fix, the bug is in CalciteRelNodeVisitor.java (lines 265-321). The code does:

List<String> namedGroups = RegexCommonUtils.getNamedGroupCandidates(patternStr);
// namedGroups = ["user", "domain", "tld"]for (inti = 0; i < namedGroups.size(); i++) {
extractCall = PPLFuncImpTable.INSTANCE.resolve(
context.rexBuilder,
BuiltinFunctionName.REX_EXTRACT,
fieldRex,
context.rexBuilder.makeLiteral(patternStr),
context.relBuilder.literal(i + 1)); // ← WRONG: Assumes sequential named groups// ...
}

The code assumes named groups are at indices 1, 2, 3, ... but the actual indices are 1, 3, 5, ... due to the unnamed nested groups.

With the above buggy logic:

  • REX_EXTRACT(field, pattern, 1) → Gets Group 1 (?<user>...) = "amberduke" → CORRECT
  • REX_EXTRACT(field, pattern, 2) → Gets Group 2 (amber|hattie) = "amber" → WRONG
  • REX_EXTRACT(field, pattern, 3) → Gets Group 3 (?<domain>...) = "pyrami" → WRONG

The second and third extractions are off by one group because they hit the unnamed nested groups.

LogicalProject(
user=[REX_EXTRACT($7, '(?<user>(amber|hattie))[a-z]*@(?<domain>(pyrami|netagy))\.(?<tld>(com|org))', 1)],
domain=[REX_EXTRACT($7, '(?<user>(amber|hattie))[a-z]*@(?<domain>(pyrami|netagy))\.(?<tld>(com|org))', 2)], -- Wrong!
tld=[REX_EXTRACT($7, '(?<user>(amber|hattie))[a-z]*@(?<domain>(pyrami|netagy))\.(?<tld>(com|org))', 3)] -- Wrong!
)

Related Issues

Check List

  • New functionality includes testing.
  • New functionality has been documented.
  • New functionality has javadoc added.
  • New functionality has a user manual doc added.
  • New PPL command checklist all confirmed.
  • API changes companion pull request created.
  • Commits are signed per the DCO using --signoff or -s.
  • Public documentation issue/PR created.

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.
For more information on following Developer Certificate of Origin and signing off your commits, please check here.

matchCount++;
} else {
// If extractor returns null, it might indicate an error (like invalid group name)
// Stop processing to avoid infinite loop

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm currently thinking about adding an error handling here

Signed-off-by: Jialiang Liang <jiallian@amazon.com>
Signed-off-by: Jialiang Liang <jiallian@amazon.com>
Signed-off-by: Jialiang Liang <jiallian@amazon.com>
Signed-off-by: Jialiang Liang <jiallian@amazon.com>
Signed-off-by: Jialiang Liang <jiallian@amazon.com>
Signed-off-by: Jialiang Liang <jiallian@amazon.com>
@RyanL1997RyanL1997 changed the title [BugFix] Fix the off-by-one error for rex with nested capture groups in named groups [BugFix] Fix unexpected shift of extraction for rex with nested capture groups in named groups Oct 25, 2025
Swiddis
Swiddis previously approved these changes Oct 28, 2025
fieldRex,
context.rexBuilder.makeLiteral(patternStr),
context.relBuilder.literal(i + 1),
context.rexBuilder.makeLiteral(groupName),

@dai-chendai-chenOct 28, 2025

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I found there is a namedGroups() API in Matcher (since JDK 20?). If we can get correct index here, we don't need to modify the UDFs below? Alternatively we can move capture name -> index logic here from UDFs?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That is correct - the core issue is matching named groups to their correct indices, and Pattern.namedGroups() would be the perfect solution. However, I discovered that we're blocked by a
compatibility constraint:

  • Pattern.namedGroups() was introduced in JDK 20
  • We need backward compatibility with JDK 11/17 - for 2.19-dev

I agree that directly leveraging the Pattern.namedGroups() is the right architectural approach - we should definitely migrate to it when we fully upgrade to JDK 20+. At that point, it would be a simple one-line change in CalciteRelNodeVisitor.

Signed-off-by: Jialiang Liang <jiallian@amazon.com>
"Rex pattern must contain at least one named capture group");
}

// TODO: Once JDK 20+ is supported, consider using Pattern.namedGroups() API for more efficient

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

as for now, I added a TODO here @dai-chen

@RyanL1997
RyanL1997 merged commit 0c1ec27 into opensearch-project:mainOct 29, 2025
51 of 52 checks passed
@RyanL1997
RyanL1997 deleted the rex-extract-fix branch October 29, 2025 00:34
opensearch-trigger-botBot pushed a commit that referenced this pull request Oct 29, 2025
…ture groups in named groups (#4641)
(cherry picked from commit 0c1ec27)
Signed-off-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
RyanL1997 pushed a commit that referenced this pull request Oct 29, 2025
…ture groups in named groups (#4641) (#4692)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
expani pushed a commit to vinaykpud/sql that referenced this pull request Nov 4, 2025
sandeshkr419 added a commit to sandeshkr419/sql that referenced this pull request Dec 3, 2025
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Simeon Widdis <sawiddis@amazon.com>
Co-authored-by: Manasvini B S <manasvis@amazon.com>
Co-authored-by: opensearch-ci-bot <opensearch-infra@amazon.com>
Co-authored-by: Louis Chu <clingzhi@amazon.com>
Co-authored-by: Chen Dai <daichen@amazon.com>
Co-authored-by: Mebsina <cnoramut@gmail.com>
Co-authored-by: Yuanchun Shen <yuanchu@amazon.com>
Co-authored-by: opensearch-trigger-bot[bot] <98922864+opensearch-trigger-bot[bot]@users.noreply.github.com>
Co-authored-by: Kai Huang <105710027+ahkcs@users.noreply.github.com>
Co-authored-by: Peng Huo <penghuo@gmail.com>
Co-authored-by: Alexey Temnikov <alexey.temnikov@improving.com>
Co-authored-by: Riley Jerger <214163063+RileyJergerAmazon@users.noreply.github.com>
Co-authored-by: Tomoyuki MORITA <moritato@amazon.com>
Co-authored-by: Lantao Jin <ltjin@amazon.com>
Co-authored-by: Songkan Tang <songkant@amazon.com>
Co-authored-by: qianheng <qianheng@amazon.com>
Co-authored-by: Simeon Widdis <sawiddis@gmail.com>
Co-authored-by: Xinyuan Lu <xinyual@amazon.com>
Co-authored-by: Jialiang Liang <jiallian@amazon.com>
Co-authored-by: Peter Zhu <zhujiaxi@amazon.com>
Co-authored-by: Vinay Krishna Pudyodu <vinkrish.neo@gmail.com>
Co-authored-by: expani <anijainc@amazon.com>
Co-authored-by: expani1729 <110471048+expani@users.noreply.github.com>
Co-authored-by: Vamsi Manohar <reddyvam@amazon.com>
Co-authored-by: ritvibhatt <53196324+ritvibhatt@users.noreply.github.com>
Co-authored-by: Xinyu Hao <75524174+ishaoxy@users.noreply.github.com>
Co-authored-by: Marc Handalian <marc.handalian@gmail.com>
Co-authored-by: Marc Handalian <handalm@amazon.com>
Fix join type ambiguous issue when specify the join type with sql-like join criteria (opensearch-project#4474)
Fix issue 4441 (opensearch-project#4449)
Fix missing keywordsCanBeId (opensearch-project#4491)
Fix the bug of explicit makeNullLiteral for UDT fields (opensearch-project#4475)
Fix mapping after aggregation push down (opensearch-project#4500)
Fix percentile bug (opensearch-project#4539)
Fix JsonExtractAllFunctionIT failure (opensearch-project#4556)
Fix sort push down into agg after project already pushed (opensearch-project#4546)
Fix push down failure for min/max on derived field (opensearch-project#4572)
Fix compile issue in main (opensearch-project#4608)
Fix filter parsing failure on date fields with non-default format (opensearch-project#4616)
Fix bin nested fields issue (opensearch-project#4606)
Fix: Support Alias Fields in MIN, MAX, FIRST, LAST, and TAKE Aggregations (opensearch-project#4621)
fix rename issue (opensearch-project#4670)
Fixes for `Multisearch` and `Append` command (opensearch-project#4512)
Fix asc/desc keyword behavior for sort command (opensearch-project#4651)
Fix] Fix unexpected shift of extraction for `rex` with nested capture groups in named groups (opensearch-project#4641)
Fix CVE-2025-48924 (opensearch-project#4665)
Fix sub-fields accessing of generated structs (opensearch-project#4683)
Fix] Incorrect Field Index Mapping in AVG to SUM/COUNT Conversion (opensearch-project#15)
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

backport 2.19-devbugSomething isn't workingbugFixPPLPiped processing language

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[BUG] rex command off-by-one error with nested capture groups in named groups

3 participants

@RyanL1997@Swiddis@dai-chen
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[BugFix] Fix unexpected shift of extraction for rex with nested capture groups in named groups - #4641

Merged
RyanL1997 merged 7 commits into
opensearch-project:mainfrom
RyanL1997:rex-extract-fix
Oct 29, 2025
Merged

[BugFix] Fix unexpected shift of extraction for rex with nested capture groups in named groups #4641
RyanL1997 merged 7 commits into
opensearch-project:mainfrom
RyanL1997:rex-extract-fix

Conversation

@RyanL1997

@RyanL1997RyanL1997 commented Oct 22, 2025

Copy link
Copy Markdown
Collaborator

Description

Fix unexpected shift of extraction for rex with nested capture groups in named groups

The rex command in PPL had a critical bug when using named capture groups that contained nested unnamed groups. This caused extracted field values to shift by one position, producing incorrect results.

  • Root Cause: Code used sequential indices 1, 2, 3... but nested groups create non-sequential indices 1, 3, 5...
  • Solution: Bypass index calculation entirely by using Java's native named group extraction (matcher.group(groupName))

Example of the Bug

Query:

curl -X POST "localhost:9200/_plugins/_ppl" \
-H "Content-Type: application/json" \
-d '{ "query": "source=accounts | rex field=email \"(?<user>(amber|hattie|nanette)[a-z]*)@(?<domain>(pyrami|netagy|quility))\\.(?<tld>(com|org))\" | fields user, domain, tld | head 1" }'

Expected Result (correct):

["amberduke", "pyrami", "com"]

Actual Result (wrong):

["amberduke", "amber", "pyrami"]

Root Cause

When Java's regex engine processes the pattern (?<user>(amber|hattie))[a-z]*, it assigns group numbers to ALL capture groups (named and unnamed):

Pattern: (?<user>(amber|hattie))[a-z]*@(?<domain>(pyrami|netagy))\.(?<tld>(com|org))

Group Assignment:

  • Group 0: Entire match
  • Group 1: (?<user>...) ← Named group "user"
  • Group 2: (amber|hattie) ← Unnamed nested group
  • Group 3: (?<domain>...) ← Named group "domain"
  • Group 4: (pyrami|netagy) ← Unnamed nested group
  • Group 5: (?<tld>...) ← Named group "tld"
  • Group 6: (com|org) ← Unnamed nested group

Before the fix, the bug is in CalciteRelNodeVisitor.java (lines 265-321). The code does:

List<String> namedGroups = RegexCommonUtils.getNamedGroupCandidates(patternStr);
// namedGroups = ["user", "domain", "tld"]for (inti = 0; i < namedGroups.size(); i++) {
extractCall = PPLFuncImpTable.INSTANCE.resolve(
context.rexBuilder,
BuiltinFunctionName.REX_EXTRACT,
fieldRex,
context.rexBuilder.makeLiteral(patternStr),
context.relBuilder.literal(i + 1)); // ← WRONG: Assumes sequential named groups// ...
}

The code assumes named groups are at indices 1, 2, 3, ... but the actual indices are 1, 3, 5, ... due to the unnamed nested groups.

With the above buggy logic:

  • REX_EXTRACT(field, pattern, 1) → Gets Group 1 (?<user>...) = "amberduke" → CORRECT
  • REX_EXTRACT(field, pattern, 2) → Gets Group 2 (amber|hattie) = "amber" → WRONG
  • REX_EXTRACT(field, pattern, 3) → Gets Group 3 (?<domain>...) = "pyrami" → WRONG

The second and third extractions are off by one group because they hit the unnamed nested groups.

LogicalProject(
user=[REX_EXTRACT($7, '(?<user>(amber|hattie))[a-z]*@(?<domain>(pyrami|netagy))\.(?<tld>(com|org))', 1)],
domain=[REX_EXTRACT($7, '(?<user>(amber|hattie))[a-z]*@(?<domain>(pyrami|netagy))\.(?<tld>(com|org))', 2)], -- Wrong!
tld=[REX_EXTRACT($7, '(?<user>(amber|hattie))[a-z]*@(?<domain>(pyrami|netagy))\.(?<tld>(com|org))', 3)] -- Wrong!
)

Related Issues

Check List

  • New functionality includes testing.
  • New functionality has been documented.
  • New functionality has javadoc added.
  • New functionality has a user manual doc added.
  • New PPL command checklist all confirmed.
  • API changes companion pull request created.
  • Commits are signed per the DCO using --signoff or -s.
  • Public documentation issue/PR created.

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.
For more information on following Developer Certificate of Origin and signing off your commits, please check here.

matchCount++;
} else {
// If extractor returns null, it might indicate an error (like invalid group name)
// Stop processing to avoid infinite loop

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm currently thinking about adding an error handling here

Signed-off-by: Jialiang Liang <jiallian@amazon.com>
Signed-off-by: Jialiang Liang <jiallian@amazon.com>
Signed-off-by: Jialiang Liang <jiallian@amazon.com>
Signed-off-by: Jialiang Liang <jiallian@amazon.com>
Signed-off-by: Jialiang Liang <jiallian@amazon.com>
Signed-off-by: Jialiang Liang <jiallian@amazon.com>
@RyanL1997RyanL1997 changed the title [BugFix] Fix the off-by-one error for rex with nested capture groups in named groups [BugFix] Fix unexpected shift of extraction for rex with nested capture groups in named groups Oct 25, 2025
Swiddis
Swiddis previously approved these changes Oct 28, 2025
fieldRex,
context.rexBuilder.makeLiteral(patternStr),
context.relBuilder.literal(i + 1),
context.rexBuilder.makeLiteral(groupName),

@dai-chendai-chenOct 28, 2025

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I found there is a namedGroups() API in Matcher (since JDK 20?). If we can get correct index here, we don't need to modify the UDFs below? Alternatively we can move capture name -> index logic here from UDFs?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That is correct - the core issue is matching named groups to their correct indices, and Pattern.namedGroups() would be the perfect solution. However, I discovered that we're blocked by a
compatibility constraint:

  • Pattern.namedGroups() was introduced in JDK 20
  • We need backward compatibility with JDK 11/17 - for 2.19-dev

I agree that directly leveraging the Pattern.namedGroups() is the right architectural approach - we should definitely migrate to it when we fully upgrade to JDK 20+. At that point, it would be a simple one-line change in CalciteRelNodeVisitor.

Signed-off-by: Jialiang Liang <jiallian@amazon.com>
"Rex pattern must contain at least one named capture group");
}

// TODO: Once JDK 20+ is supported, consider using Pattern.namedGroups() API for more efficient

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

as for now, I added a TODO here @dai-chen

@RyanL1997
RyanL1997 merged commit 0c1ec27 into opensearch-project:mainOct 29, 2025
51 of 52 checks passed
@RyanL1997
RyanL1997 deleted the rex-extract-fix branch October 29, 2025 00:34
opensearch-trigger-botBot pushed a commit that referenced this pull request Oct 29, 2025
…ture groups in named groups (#4641)
(cherry picked from commit 0c1ec27)
Signed-off-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
RyanL1997 pushed a commit that referenced this pull request Oct 29, 2025
…ture groups in named groups (#4641) (#4692)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
expani pushed a commit to vinaykpud/sql that referenced this pull request Nov 4, 2025
sandeshkr419 added a commit to sandeshkr419/sql that referenced this pull request Dec 3, 2025
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Simeon Widdis <sawiddis@amazon.com>
Co-authored-by: Manasvini B S <manasvis@amazon.com>
Co-authored-by: opensearch-ci-bot <opensearch-infra@amazon.com>
Co-authored-by: Louis Chu <clingzhi@amazon.com>
Co-authored-by: Chen Dai <daichen@amazon.com>
Co-authored-by: Mebsina <cnoramut@gmail.com>
Co-authored-by: Yuanchun Shen <yuanchu@amazon.com>
Co-authored-by: opensearch-trigger-bot[bot] <98922864+opensearch-trigger-bot[bot]@users.noreply.github.com>
Co-authored-by: Kai Huang <105710027+ahkcs@users.noreply.github.com>
Co-authored-by: Peng Huo <penghuo@gmail.com>
Co-authored-by: Alexey Temnikov <alexey.temnikov@improving.com>
Co-authored-by: Riley Jerger <214163063+RileyJergerAmazon@users.noreply.github.com>
Co-authored-by: Tomoyuki MORITA <moritato@amazon.com>
Co-authored-by: Lantao Jin <ltjin@amazon.com>
Co-authored-by: Songkan Tang <songkant@amazon.com>
Co-authored-by: qianheng <qianheng@amazon.com>
Co-authored-by: Simeon Widdis <sawiddis@gmail.com>
Co-authored-by: Xinyuan Lu <xinyual@amazon.com>
Co-authored-by: Jialiang Liang <jiallian@amazon.com>
Co-authored-by: Peter Zhu <zhujiaxi@amazon.com>
Co-authored-by: Vinay Krishna Pudyodu <vinkrish.neo@gmail.com>
Co-authored-by: expani <anijainc@amazon.com>
Co-authored-by: expani1729 <110471048+expani@users.noreply.github.com>
Co-authored-by: Vamsi Manohar <reddyvam@amazon.com>
Co-authored-by: ritvibhatt <53196324+ritvibhatt@users.noreply.github.com>
Co-authored-by: Xinyu Hao <75524174+ishaoxy@users.noreply.github.com>
Co-authored-by: Marc Handalian <marc.handalian@gmail.com>
Co-authored-by: Marc Handalian <handalm@amazon.com>
Fix join type ambiguous issue when specify the join type with sql-like join criteria (opensearch-project#4474)
Fix issue 4441 (opensearch-project#4449)
Fix missing keywordsCanBeId (opensearch-project#4491)
Fix the bug of explicit makeNullLiteral for UDT fields (opensearch-project#4475)
Fix mapping after aggregation push down (opensearch-project#4500)
Fix percentile bug (opensearch-project#4539)
Fix JsonExtractAllFunctionIT failure (opensearch-project#4556)
Fix sort push down into agg after project already pushed (opensearch-project#4546)
Fix push down failure for min/max on derived field (opensearch-project#4572)
Fix compile issue in main (opensearch-project#4608)
Fix filter parsing failure on date fields with non-default format (opensearch-project#4616)
Fix bin nested fields issue (opensearch-project#4606)
Fix: Support Alias Fields in MIN, MAX, FIRST, LAST, and TAKE Aggregations (opensearch-project#4621)
fix rename issue (opensearch-project#4670)
Fixes for `Multisearch` and `Append` command (opensearch-project#4512)
Fix asc/desc keyword behavior for sort command (opensearch-project#4651)
Fix] Fix unexpected shift of extraction for `rex` with nested capture groups in named groups (opensearch-project#4641)
Fix CVE-2025-48924 (opensearch-project#4665)
Fix sub-fields accessing of generated structs (opensearch-project#4683)
Fix] Incorrect Field Index Mapping in AVG to SUM/COUNT Conversion (opensearch-project#15)
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

backport 2.19-devbugSomething isn't workingbugFixPPLPiped processing language

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[BUG] rex command off-by-one error with nested capture groups in named groups

3 participants

@RyanL1997@Swiddis@dai-chen
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[BugFix] Fix unexpected shift of extraction for rex with nested capture groups in named groups - #4641

Merged
RyanL1997 merged 7 commits into
opensearch-project:mainfrom
RyanL1997:rex-extract-fix
Oct 29, 2025
Merged

[BugFix] Fix unexpected shift of extraction for rex with nested capture groups in named groups #4641
RyanL1997 merged 7 commits into
opensearch-project:mainfrom
RyanL1997:rex-extract-fix

Conversation

@RyanL1997

@RyanL1997RyanL1997 commented Oct 22, 2025

Copy link
Copy Markdown
Collaborator

Description

Fix unexpected shift of extraction for rex with nested capture groups in named groups

The rex command in PPL had a critical bug when using named capture groups that contained nested unnamed groups. This caused extracted field values to shift by one position, producing incorrect results.

  • Root Cause: Code used sequential indices 1, 2, 3... but nested groups create non-sequential indices 1, 3, 5...
  • Solution: Bypass index calculation entirely by using Java's native named group extraction (matcher.group(groupName))

Example of the Bug

Query:

curl -X POST "localhost:9200/_plugins/_ppl" \
-H "Content-Type: application/json" \
-d '{ "query": "source=accounts | rex field=email \"(?<user>(amber|hattie|nanette)[a-z]*)@(?<domain>(pyrami|netagy|quility))\\.(?<tld>(com|org))\" | fields user, domain, tld | head 1" }'

Expected Result (correct):

["amberduke", "pyrami", "com"]

Actual Result (wrong):

["amberduke", "amber", "pyrami"]

Root Cause

When Java's regex engine processes the pattern (?<user>(amber|hattie))[a-z]*, it assigns group numbers to ALL capture groups (named and unnamed):

Pattern: (?<user>(amber|hattie))[a-z]*@(?<domain>(pyrami|netagy))\.(?<tld>(com|org))

Group Assignment:

  • Group 0: Entire match
  • Group 1: (?<user>...) ← Named group "user"
  • Group 2: (amber|hattie) ← Unnamed nested group
  • Group 3: (?<domain>...) ← Named group "domain"
  • Group 4: (pyrami|netagy) ← Unnamed nested group
  • Group 5: (?<tld>...) ← Named group "tld"
  • Group 6: (com|org) ← Unnamed nested group

Before the fix, the bug is in CalciteRelNodeVisitor.java (lines 265-321). The code does:

List<String> namedGroups = RegexCommonUtils.getNamedGroupCandidates(patternStr);
// namedGroups = ["user", "domain", "tld"]for (inti = 0; i < namedGroups.size(); i++) {
extractCall = PPLFuncImpTable.INSTANCE.resolve(
context.rexBuilder,
BuiltinFunctionName.REX_EXTRACT,
fieldRex,
context.rexBuilder.makeLiteral(patternStr),
context.relBuilder.literal(i + 1)); // ← WRONG: Assumes sequential named groups// ...
}

The code assumes named groups are at indices 1, 2, 3, ... but the actual indices are 1, 3, 5, ... due to the unnamed nested groups.

With the above buggy logic:

  • REX_EXTRACT(field, pattern, 1) → Gets Group 1 (?<user>...) = "amberduke" → CORRECT
  • REX_EXTRACT(field, pattern, 2) → Gets Group 2 (amber|hattie) = "amber" → WRONG
  • REX_EXTRACT(field, pattern, 3) → Gets Group 3 (?<domain>...) = "pyrami" → WRONG

The second and third extractions are off by one group because they hit the unnamed nested groups.

LogicalProject(
user=[REX_EXTRACT($7, '(?<user>(amber|hattie))[a-z]*@(?<domain>(pyrami|netagy))\.(?<tld>(com|org))', 1)],
domain=[REX_EXTRACT($7, '(?<user>(amber|hattie))[a-z]*@(?<domain>(pyrami|netagy))\.(?<tld>(com|org))', 2)], -- Wrong!
tld=[REX_EXTRACT($7, '(?<user>(amber|hattie))[a-z]*@(?<domain>(pyrami|netagy))\.(?<tld>(com|org))', 3)] -- Wrong!
)

Related Issues

Check List

  • New functionality includes testing.
  • New functionality has been documented.
  • New functionality has javadoc added.
  • New functionality has a user manual doc added.
  • New PPL command checklist all confirmed.
  • API changes companion pull request created.
  • Commits are signed per the DCO using --signoff or -s.
  • Public documentation issue/PR created.

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.
For more information on following Developer Certificate of Origin and signing off your commits, please check here.

matchCount++;
} else {
// If extractor returns null, it might indicate an error (like invalid group name)
// Stop processing to avoid infinite loop

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm currently thinking about adding an error handling here

Signed-off-by: Jialiang Liang <jiallian@amazon.com>
Signed-off-by: Jialiang Liang <jiallian@amazon.com>
Signed-off-by: Jialiang Liang <jiallian@amazon.com>
Signed-off-by: Jialiang Liang <jiallian@amazon.com>
Signed-off-by: Jialiang Liang <jiallian@amazon.com>
Signed-off-by: Jialiang Liang <jiallian@amazon.com>
@RyanL1997RyanL1997 changed the title [BugFix] Fix the off-by-one error for rex with nested capture groups in named groups [BugFix] Fix unexpected shift of extraction for rex with nested capture groups in named groups Oct 25, 2025
Swiddis
Swiddis previously approved these changes Oct 28, 2025
fieldRex,
context.rexBuilder.makeLiteral(patternStr),
context.relBuilder.literal(i + 1),
context.rexBuilder.makeLiteral(groupName),

@dai-chendai-chenOct 28, 2025

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I found there is a namedGroups() API in Matcher (since JDK 20?). If we can get correct index here, we don't need to modify the UDFs below? Alternatively we can move capture name -> index logic here from UDFs?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That is correct - the core issue is matching named groups to their correct indices, and Pattern.namedGroups() would be the perfect solution. However, I discovered that we're blocked by a
compatibility constraint:

  • Pattern.namedGroups() was introduced in JDK 20
  • We need backward compatibility with JDK 11/17 - for 2.19-dev

I agree that directly leveraging the Pattern.namedGroups() is the right architectural approach - we should definitely migrate to it when we fully upgrade to JDK 20+. At that point, it would be a simple one-line change in CalciteRelNodeVisitor.

Signed-off-by: Jialiang Liang <jiallian@amazon.com>
"Rex pattern must contain at least one named capture group");
}

// TODO: Once JDK 20+ is supported, consider using Pattern.namedGroups() API for more efficient

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

as for now, I added a TODO here @dai-chen

@RyanL1997
RyanL1997 merged commit 0c1ec27 into opensearch-project:mainOct 29, 2025
51 of 52 checks passed
@RyanL1997
RyanL1997 deleted the rex-extract-fix branch October 29, 2025 00:34
opensearch-trigger-botBot pushed a commit that referenced this pull request Oct 29, 2025
…ture groups in named groups (#4641)
(cherry picked from commit 0c1ec27)
Signed-off-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
RyanL1997 pushed a commit that referenced this pull request Oct 29, 2025
…ture groups in named groups (#4641) (#4692)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
expani pushed a commit to vinaykpud/sql that referenced this pull request Nov 4, 2025
sandeshkr419 added a commit to sandeshkr419/sql that referenced this pull request Dec 3, 2025
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Simeon Widdis <sawiddis@amazon.com>
Co-authored-by: Manasvini B S <manasvis@amazon.com>
Co-authored-by: opensearch-ci-bot <opensearch-infra@amazon.com>
Co-authored-by: Louis Chu <clingzhi@amazon.com>
Co-authored-by: Chen Dai <daichen@amazon.com>
Co-authored-by: Mebsina <cnoramut@gmail.com>
Co-authored-by: Yuanchun Shen <yuanchu@amazon.com>
Co-authored-by: opensearch-trigger-bot[bot] <98922864+opensearch-trigger-bot[bot]@users.noreply.github.com>
Co-authored-by: Kai Huang <105710027+ahkcs@users.noreply.github.com>
Co-authored-by: Peng Huo <penghuo@gmail.com>
Co-authored-by: Alexey Temnikov <alexey.temnikov@improving.com>
Co-authored-by: Riley Jerger <214163063+RileyJergerAmazon@users.noreply.github.com>
Co-authored-by: Tomoyuki MORITA <moritato@amazon.com>
Co-authored-by: Lantao Jin <ltjin@amazon.com>
Co-authored-by: Songkan Tang <songkant@amazon.com>
Co-authored-by: qianheng <qianheng@amazon.com>
Co-authored-by: Simeon Widdis <sawiddis@gmail.com>
Co-authored-by: Xinyuan Lu <xinyual@amazon.com>
Co-authored-by: Jialiang Liang <jiallian@amazon.com>
Co-authored-by: Peter Zhu <zhujiaxi@amazon.com>
Co-authored-by: Vinay Krishna Pudyodu <vinkrish.neo@gmail.com>
Co-authored-by: expani <anijainc@amazon.com>
Co-authored-by: expani1729 <110471048+expani@users.noreply.github.com>
Co-authored-by: Vamsi Manohar <reddyvam@amazon.com>
Co-authored-by: ritvibhatt <53196324+ritvibhatt@users.noreply.github.com>
Co-authored-by: Xinyu Hao <75524174+ishaoxy@users.noreply.github.com>
Co-authored-by: Marc Handalian <marc.handalian@gmail.com>
Co-authored-by: Marc Handalian <handalm@amazon.com>
Fix join type ambiguous issue when specify the join type with sql-like join criteria (opensearch-project#4474)
Fix issue 4441 (opensearch-project#4449)
Fix missing keywordsCanBeId (opensearch-project#4491)
Fix the bug of explicit makeNullLiteral for UDT fields (opensearch-project#4475)
Fix mapping after aggregation push down (opensearch-project#4500)
Fix percentile bug (opensearch-project#4539)
Fix JsonExtractAllFunctionIT failure (opensearch-project#4556)
Fix sort push down into agg after project already pushed (opensearch-project#4546)
Fix push down failure for min/max on derived field (opensearch-project#4572)
Fix compile issue in main (opensearch-project#4608)
Fix filter parsing failure on date fields with non-default format (opensearch-project#4616)
Fix bin nested fields issue (opensearch-project#4606)
Fix: Support Alias Fields in MIN, MAX, FIRST, LAST, and TAKE Aggregations (opensearch-project#4621)
fix rename issue (opensearch-project#4670)
Fixes for `Multisearch` and `Append` command (opensearch-project#4512)
Fix asc/desc keyword behavior for sort command (opensearch-project#4651)
Fix] Fix unexpected shift of extraction for `rex` with nested capture groups in named groups (opensearch-project#4641)
Fix CVE-2025-48924 (opensearch-project#4665)
Fix sub-fields accessing of generated structs (opensearch-project#4683)
Fix] Incorrect Field Index Mapping in AVG to SUM/COUNT Conversion (opensearch-project#15)
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

backport 2.19-devbugSomething isn't workingbugFixPPLPiped processing language

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[BUG] rex command off-by-one error with nested capture groups in named groups

3 participants

@RyanL1997@Swiddis@dai-chen
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

[BugFix] Fix unexpected shift of extraction for rex with nested capture groups in named groups - #4641

Merged
RyanL1997 merged 7 commits into
opensearch-project:mainfrom
RyanL1997:rex-extract-fix
Oct 29, 2025
Merged

[BugFix] Fix unexpected shift of extraction for rex with nested capture groups in named groups #4641
RyanL1997 merged 7 commits into
opensearch-project:mainfrom
RyanL1997:rex-extract-fix

Conversation

@RyanL1997

@RyanL1997RyanL1997 commented Oct 22, 2025

Copy link
Copy Markdown
Collaborator

Description

Fix unexpected shift of extraction for rex with nested capture groups in named groups

The rex command in PPL had a critical bug when using named capture groups that contained nested unnamed groups. This caused extracted field values to shift by one position, producing incorrect results.

  • Root Cause: Code used sequential indices 1, 2, 3... but nested groups create non-sequential indices 1, 3, 5...
  • Solution: Bypass index calculation entirely by using Java's native named group extraction (matcher.group(groupName))

Example of the Bug

Query:

curl -X POST "localhost:9200/_plugins/_ppl" \
-H "Content-Type: application/json" \
-d '{ "query": "source=accounts | rex field=email \"(?<user>(amber|hattie|nanette)[a-z]*)@(?<domain>(pyrami|netagy|quility))\\.(?<tld>(com|org))\" | fields user, domain, tld | head 1" }'

Expected Result (correct):

["amberduke", "pyrami", "com"]

Actual Result (wrong):

["amberduke", "amber", "pyrami"]

Root Cause

When Java's regex engine processes the pattern (?<user>(amber|hattie))[a-z]*, it assigns group numbers to ALL capture groups (named and unnamed):

Pattern: (?<user>(amber|hattie))[a-z]*@(?<domain>(pyrami|netagy))\.(?<tld>(com|org))

Group Assignment:

  • Group 0: Entire match
  • Group 1: (?<user>...) ← Named group "user"
  • Group 2: (amber|hattie) ← Unnamed nested group
  • Group 3: (?<domain>...) ← Named group "domain"
  • Group 4: (pyrami|netagy) ← Unnamed nested group
  • Group 5: (?<tld>...) ← Named group "tld"
  • Group 6: (com|org) ← Unnamed nested group

Before the fix, the bug is in CalciteRelNodeVisitor.java (lines 265-321). The code does:

List<String> namedGroups = RegexCommonUtils.getNamedGroupCandidates(patternStr);
// namedGroups = ["user", "domain", "tld"]for (inti = 0; i < namedGroups.size(); i++) {
extractCall = PPLFuncImpTable.INSTANCE.resolve(
context.rexBuilder,
BuiltinFunctionName.REX_EXTRACT,
fieldRex,
context.rexBuilder.makeLiteral(patternStr),
context.relBuilder.literal(i + 1)); // ← WRONG: Assumes sequential named groups// ...
}

The code assumes named groups are at indices 1, 2, 3, ... but the actual indices are 1, 3, 5, ... due to the unnamed nested groups.

With the above buggy logic:

  • REX_EXTRACT(field, pattern, 1) → Gets Group 1 (?<user>...) = "amberduke" → CORRECT
  • REX_EXTRACT(field, pattern, 2) → Gets Group 2 (amber|hattie) = "amber" → WRONG
  • REX_EXTRACT(field, pattern, 3) → Gets Group 3 (?<domain>...) = "pyrami" → WRONG

The second and third extractions are off by one group because they hit the unnamed nested groups.

LogicalProject(
user=[REX_EXTRACT($7, '(?<user>(amber|hattie))[a-z]*@(?<domain>(pyrami|netagy))\.(?<tld>(com|org))', 1)],
domain=[REX_EXTRACT($7, '(?<user>(amber|hattie))[a-z]*@(?<domain>(pyrami|netagy))\.(?<tld>(com|org))', 2)], -- Wrong!
tld=[REX_EXTRACT($7, '(?<user>(amber|hattie))[a-z]*@(?<domain>(pyrami|netagy))\.(?<tld>(com|org))', 3)] -- Wrong!
)

Related Issues

Check List

  • New functionality includes testing.
  • New functionality has been documented.
  • New functionality has javadoc added.
  • New functionality has a user manual doc added.
  • New PPL command checklist all confirmed.
  • API changes companion pull request created.
  • Commits are signed per the DCO using --signoff or -s.
  • Public documentation issue/PR created.

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.
For more information on following Developer Certificate of Origin and signing off your commits, please check here.

matchCount++;
} else {
// If extractor returns null, it might indicate an error (like invalid group name)
// Stop processing to avoid infinite loop

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm currently thinking about adding an error handling here

Signed-off-by: Jialiang Liang <jiallian@amazon.com>
Signed-off-by: Jialiang Liang <jiallian@amazon.com>
Signed-off-by: Jialiang Liang <jiallian@amazon.com>
Signed-off-by: Jialiang Liang <jiallian@amazon.com>
Signed-off-by: Jialiang Liang <jiallian@amazon.com>
Signed-off-by: Jialiang Liang <jiallian@amazon.com>
@RyanL1997RyanL1997 changed the title [BugFix] Fix the off-by-one error for rex with nested capture groups in named groups [BugFix] Fix unexpected shift of extraction for rex with nested capture groups in named groups Oct 25, 2025
Swiddis
Swiddis previously approved these changes Oct 28, 2025
fieldRex,
context.rexBuilder.makeLiteral(patternStr),
context.relBuilder.literal(i + 1),
context.rexBuilder.makeLiteral(groupName),

@dai-chendai-chenOct 28, 2025

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I found there is a namedGroups() API in Matcher (since JDK 20?). If we can get correct index here, we don't need to modify the UDFs below? Alternatively we can move capture name -> index logic here from UDFs?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That is correct - the core issue is matching named groups to their correct indices, and Pattern.namedGroups() would be the perfect solution. However, I discovered that we're blocked by a
compatibility constraint:

  • Pattern.namedGroups() was introduced in JDK 20
  • We need backward compatibility with JDK 11/17 - for 2.19-dev

I agree that directly leveraging the Pattern.namedGroups() is the right architectural approach - we should definitely migrate to it when we fully upgrade to JDK 20+. At that point, it would be a simple one-line change in CalciteRelNodeVisitor.

Signed-off-by: Jialiang Liang <jiallian@amazon.com>
"Rex pattern must contain at least one named capture group");
}

// TODO: Once JDK 20+ is supported, consider using Pattern.namedGroups() API for more efficient

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

as for now, I added a TODO here @dai-chen

@RyanL1997
RyanL1997 merged commit 0c1ec27 into opensearch-project:mainOct 29, 2025
51 of 52 checks passed
@RyanL1997
RyanL1997 deleted the rex-extract-fix branch October 29, 2025 00:34
opensearch-trigger-botBot pushed a commit that referenced this pull request Oct 29, 2025
…ture groups in named groups (#4641)
(cherry picked from commit 0c1ec27)
Signed-off-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
RyanL1997 pushed a commit that referenced this pull request Oct 29, 2025
…ture groups in named groups (#4641) (#4692)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
expani pushed a commit to vinaykpud/sql that referenced this pull request Nov 4, 2025
sandeshkr419 added a commit to sandeshkr419/sql that referenced this pull request Dec 3, 2025
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Simeon Widdis <sawiddis@amazon.com>
Co-authored-by: Manasvini B S <manasvis@amazon.com>
Co-authored-by: opensearch-ci-bot <opensearch-infra@amazon.com>
Co-authored-by: Louis Chu <clingzhi@amazon.com>
Co-authored-by: Chen Dai <daichen@amazon.com>
Co-authored-by: Mebsina <cnoramut@gmail.com>
Co-authored-by: Yuanchun Shen <yuanchu@amazon.com>
Co-authored-by: opensearch-trigger-bot[bot] <98922864+opensearch-trigger-bot[bot]@users.noreply.github.com>
Co-authored-by: Kai Huang <105710027+ahkcs@users.noreply.github.com>
Co-authored-by: Peng Huo <penghuo@gmail.com>
Co-authored-by: Alexey Temnikov <alexey.temnikov@improving.com>
Co-authored-by: Riley Jerger <214163063+RileyJergerAmazon@users.noreply.github.com>
Co-authored-by: Tomoyuki MORITA <moritato@amazon.com>
Co-authored-by: Lantao Jin <ltjin@amazon.com>
Co-authored-by: Songkan Tang <songkant@amazon.com>
Co-authored-by: qianheng <qianheng@amazon.com>
Co-authored-by: Simeon Widdis <sawiddis@gmail.com>
Co-authored-by: Xinyuan Lu <xinyual@amazon.com>
Co-authored-by: Jialiang Liang <jiallian@amazon.com>
Co-authored-by: Peter Zhu <zhujiaxi@amazon.com>
Co-authored-by: Vinay Krishna Pudyodu <vinkrish.neo@gmail.com>
Co-authored-by: expani <anijainc@amazon.com>
Co-authored-by: expani1729 <110471048+expani@users.noreply.github.com>
Co-authored-by: Vamsi Manohar <reddyvam@amazon.com>
Co-authored-by: ritvibhatt <53196324+ritvibhatt@users.noreply.github.com>
Co-authored-by: Xinyu Hao <75524174+ishaoxy@users.noreply.github.com>
Co-authored-by: Marc Handalian <marc.handalian@gmail.com>
Co-authored-by: Marc Handalian <handalm@amazon.com>
Fix join type ambiguous issue when specify the join type with sql-like join criteria (opensearch-project#4474)
Fix issue 4441 (opensearch-project#4449)
Fix missing keywordsCanBeId (opensearch-project#4491)
Fix the bug of explicit makeNullLiteral for UDT fields (opensearch-project#4475)
Fix mapping after aggregation push down (opensearch-project#4500)
Fix percentile bug (opensearch-project#4539)
Fix JsonExtractAllFunctionIT failure (opensearch-project#4556)
Fix sort push down into agg after project already pushed (opensearch-project#4546)
Fix push down failure for min/max on derived field (opensearch-project#4572)
Fix compile issue in main (opensearch-project#4608)
Fix filter parsing failure on date fields with non-default format (opensearch-project#4616)
Fix bin nested fields issue (opensearch-project#4606)
Fix: Support Alias Fields in MIN, MAX, FIRST, LAST, and TAKE Aggregations (opensearch-project#4621)
fix rename issue (opensearch-project#4670)
Fixes for `Multisearch` and `Append` command (opensearch-project#4512)
Fix asc/desc keyword behavior for sort command (opensearch-project#4651)
Fix] Fix unexpected shift of extraction for `rex` with nested capture groups in named groups (opensearch-project#4641)
Fix CVE-2025-48924 (opensearch-project#4665)
Fix sub-fields accessing of generated structs (opensearch-project#4683)
Fix] Incorrect Field Index Mapping in AVG to SUM/COUNT Conversion (opensearch-project#15)
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

backport 2.19-devbugSomething isn't workingbugFixPPLPiped processing language

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[BUG] rex command off-by-one error with nested capture groups in named groups

3 participants

@RyanL1997@Swiddis@dai-chen
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[BugFix] Fix unexpected shift of extraction for rex with nested capture groups in named groups - #4641

Merged
RyanL1997 merged 7 commits into
opensearch-project:mainfrom
RyanL1997:rex-extract-fix
Oct 29, 2025
Merged

[BugFix] Fix unexpected shift of extraction for rex with nested capture groups in named groups #4641
RyanL1997 merged 7 commits into
opensearch-project:mainfrom
RyanL1997:rex-extract-fix

Conversation

@RyanL1997

@RyanL1997RyanL1997 commented Oct 22, 2025

Copy link
Copy Markdown
Collaborator

Description

Fix unexpected shift of extraction for rex with nested capture groups in named groups

The rex command in PPL had a critical bug when using named capture groups that contained nested unnamed groups. This caused extracted field values to shift by one position, producing incorrect results.

  • Root Cause: Code used sequential indices 1, 2, 3... but nested groups create non-sequential indices 1, 3, 5...
  • Solution: Bypass index calculation entirely by using Java's native named group extraction (matcher.group(groupName))

Example of the Bug

Query:

curl -X POST "localhost:9200/_plugins/_ppl" \
-H "Content-Type: application/json" \
-d '{ "query": "source=accounts | rex field=email \"(?<user>(amber|hattie|nanette)[a-z]*)@(?<domain>(pyrami|netagy|quility))\\.(?<tld>(com|org))\" | fields user, domain, tld | head 1" }'

Expected Result (correct):

["amberduke", "pyrami", "com"]

Actual Result (wrong):

["amberduke", "amber", "pyrami"]

Root Cause

When Java's regex engine processes the pattern (?<user>(amber|hattie))[a-z]*, it assigns group numbers to ALL capture groups (named and unnamed):

Pattern: (?<user>(amber|hattie))[a-z]*@(?<domain>(pyrami|netagy))\.(?<tld>(com|org))

Group Assignment:

  • Group 0: Entire match
  • Group 1: (?<user>...) ← Named group "user"
  • Group 2: (amber|hattie) ← Unnamed nested group
  • Group 3: (?<domain>...) ← Named group "domain"
  • Group 4: (pyrami|netagy) ← Unnamed nested group
  • Group 5: (?<tld>...) ← Named group "tld"
  • Group 6: (com|org) ← Unnamed nested group

Before the fix, the bug is in CalciteRelNodeVisitor.java (lines 265-321). The code does:

List<String> namedGroups = RegexCommonUtils.getNamedGroupCandidates(patternStr);
// namedGroups = ["user", "domain", "tld"]for (inti = 0; i < namedGroups.size(); i++) {
extractCall = PPLFuncImpTable.INSTANCE.resolve(
context.rexBuilder,
BuiltinFunctionName.REX_EXTRACT,
fieldRex,
context.rexBuilder.makeLiteral(patternStr),
context.relBuilder.literal(i + 1)); // ← WRONG: Assumes sequential named groups// ...
}

The code assumes named groups are at indices 1, 2, 3, ... but the actual indices are 1, 3, 5, ... due to the unnamed nested groups.

With the above buggy logic:

  • REX_EXTRACT(field, pattern, 1) → Gets Group 1 (?<user>...) = "amberduke" → CORRECT
  • REX_EXTRACT(field, pattern, 2) → Gets Group 2 (amber|hattie) = "amber" → WRONG
  • REX_EXTRACT(field, pattern, 3) → Gets Group 3 (?<domain>...) = "pyrami" → WRONG

The second and third extractions are off by one group because they hit the unnamed nested groups.

LogicalProject(
user=[REX_EXTRACT($7, '(?<user>(amber|hattie))[a-z]*@(?<domain>(pyrami|netagy))\.(?<tld>(com|org))', 1)],
domain=[REX_EXTRACT($7, '(?<user>(amber|hattie))[a-z]*@(?<domain>(pyrami|netagy))\.(?<tld>(com|org))', 2)], -- Wrong!
tld=[REX_EXTRACT($7, '(?<user>(amber|hattie))[a-z]*@(?<domain>(pyrami|netagy))\.(?<tld>(com|org))', 3)] -- Wrong!
)

Related Issues

Check List

  • New functionality includes testing.
  • New functionality has been documented.
  • New functionality has javadoc added.
  • New functionality has a user manual doc added.
  • New PPL command checklist all confirmed.
  • API changes companion pull request created.
  • Commits are signed per the DCO using --signoff or -s.
  • Public documentation issue/PR created.

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.
For more information on following Developer Certificate of Origin and signing off your commits, please check here.

matchCount++;
} else {
// If extractor returns null, it might indicate an error (like invalid group name)
// Stop processing to avoid infinite loop

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm currently thinking about adding an error handling here

Signed-off-by: Jialiang Liang <jiallian@amazon.com>
Signed-off-by: Jialiang Liang <jiallian@amazon.com>
Signed-off-by: Jialiang Liang <jiallian@amazon.com>
Signed-off-by: Jialiang Liang <jiallian@amazon.com>
Signed-off-by: Jialiang Liang <jiallian@amazon.com>
Signed-off-by: Jialiang Liang <jiallian@amazon.com>
@RyanL1997RyanL1997 changed the title [BugFix] Fix the off-by-one error for rex with nested capture groups in named groups [BugFix] Fix unexpected shift of extraction for rex with nested capture groups in named groups Oct 25, 2025
Swiddis
Swiddis previously approved these changes Oct 28, 2025
fieldRex,
context.rexBuilder.makeLiteral(patternStr),
context.relBuilder.literal(i + 1),
context.rexBuilder.makeLiteral(groupName),

@dai-chendai-chenOct 28, 2025

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I found there is a namedGroups() API in Matcher (since JDK 20?). If we can get correct index here, we don't need to modify the UDFs below? Alternatively we can move capture name -> index logic here from UDFs?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That is correct - the core issue is matching named groups to their correct indices, and Pattern.namedGroups() would be the perfect solution. However, I discovered that we're blocked by a
compatibility constraint:

  • Pattern.namedGroups() was introduced in JDK 20
  • We need backward compatibility with JDK 11/17 - for 2.19-dev

I agree that directly leveraging the Pattern.namedGroups() is the right architectural approach - we should definitely migrate to it when we fully upgrade to JDK 20+. At that point, it would be a simple one-line change in CalciteRelNodeVisitor.

Signed-off-by: Jialiang Liang <jiallian@amazon.com>
"Rex pattern must contain at least one named capture group");
}

// TODO: Once JDK 20+ is supported, consider using Pattern.namedGroups() API for more efficient

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

as for now, I added a TODO here @dai-chen

@RyanL1997
RyanL1997 merged commit 0c1ec27 into opensearch-project:mainOct 29, 2025
51 of 52 checks passed
@RyanL1997
RyanL1997 deleted the rex-extract-fix branch October 29, 2025 00:34
opensearch-trigger-botBot pushed a commit that referenced this pull request Oct 29, 2025
…ture groups in named groups (#4641)
(cherry picked from commit 0c1ec27)
Signed-off-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
RyanL1997 pushed a commit that referenced this pull request Oct 29, 2025
…ture groups in named groups (#4641) (#4692)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
expani pushed a commit to vinaykpud/sql that referenced this pull request Nov 4, 2025
sandeshkr419 added a commit to sandeshkr419/sql that referenced this pull request Dec 3, 2025
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Simeon Widdis <sawiddis@amazon.com>
Co-authored-by: Manasvini B S <manasvis@amazon.com>
Co-authored-by: opensearch-ci-bot <opensearch-infra@amazon.com>
Co-authored-by: Louis Chu <clingzhi@amazon.com>
Co-authored-by: Chen Dai <daichen@amazon.com>
Co-authored-by: Mebsina <cnoramut@gmail.com>
Co-authored-by: Yuanchun Shen <yuanchu@amazon.com>
Co-authored-by: opensearch-trigger-bot[bot] <98922864+opensearch-trigger-bot[bot]@users.noreply.github.com>
Co-authored-by: Kai Huang <105710027+ahkcs@users.noreply.github.com>
Co-authored-by: Peng Huo <penghuo@gmail.com>
Co-authored-by: Alexey Temnikov <alexey.temnikov@improving.com>
Co-authored-by: Riley Jerger <214163063+RileyJergerAmazon@users.noreply.github.com>
Co-authored-by: Tomoyuki MORITA <moritato@amazon.com>
Co-authored-by: Lantao Jin <ltjin@amazon.com>
Co-authored-by: Songkan Tang <songkant@amazon.com>
Co-authored-by: qianheng <qianheng@amazon.com>
Co-authored-by: Simeon Widdis <sawiddis@gmail.com>
Co-authored-by: Xinyuan Lu <xinyual@amazon.com>
Co-authored-by: Jialiang Liang <jiallian@amazon.com>
Co-authored-by: Peter Zhu <zhujiaxi@amazon.com>
Co-authored-by: Vinay Krishna Pudyodu <vinkrish.neo@gmail.com>
Co-authored-by: expani <anijainc@amazon.com>
Co-authored-by: expani1729 <110471048+expani@users.noreply.github.com>
Co-authored-by: Vamsi Manohar <reddyvam@amazon.com>
Co-authored-by: ritvibhatt <53196324+ritvibhatt@users.noreply.github.com>
Co-authored-by: Xinyu Hao <75524174+ishaoxy@users.noreply.github.com>
Co-authored-by: Marc Handalian <marc.handalian@gmail.com>
Co-authored-by: Marc Handalian <handalm@amazon.com>
Fix join type ambiguous issue when specify the join type with sql-like join criteria (opensearch-project#4474)
Fix issue 4441 (opensearch-project#4449)
Fix missing keywordsCanBeId (opensearch-project#4491)
Fix the bug of explicit makeNullLiteral for UDT fields (opensearch-project#4475)
Fix mapping after aggregation push down (opensearch-project#4500)
Fix percentile bug (opensearch-project#4539)
Fix JsonExtractAllFunctionIT failure (opensearch-project#4556)
Fix sort push down into agg after project already pushed (opensearch-project#4546)
Fix push down failure for min/max on derived field (opensearch-project#4572)
Fix compile issue in main (opensearch-project#4608)
Fix filter parsing failure on date fields with non-default format (opensearch-project#4616)
Fix bin nested fields issue (opensearch-project#4606)
Fix: Support Alias Fields in MIN, MAX, FIRST, LAST, and TAKE Aggregations (opensearch-project#4621)
fix rename issue (opensearch-project#4670)
Fixes for `Multisearch` and `Append` command (opensearch-project#4512)
Fix asc/desc keyword behavior for sort command (opensearch-project#4651)
Fix] Fix unexpected shift of extraction for `rex` with nested capture groups in named groups (opensearch-project#4641)
Fix CVE-2025-48924 (opensearch-project#4665)
Fix sub-fields accessing of generated structs (opensearch-project#4683)
Fix] Incorrect Field Index Mapping in AVG to SUM/COUNT Conversion (opensearch-project#15)
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

backport 2.19-devbugSomething isn't workingbugFixPPLPiped processing language

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[BUG] rex command off-by-one error with nested capture groups in named groups

3 participants

@RyanL1997@Swiddis@dai-chen
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[BugFix] Fix unexpected shift of extraction for rex with nested capture groups in named groups - #4641

Merged
RyanL1997 merged 7 commits into
opensearch-project:mainfrom
RyanL1997:rex-extract-fix
Oct 29, 2025
Merged

[BugFix] Fix unexpected shift of extraction for rex with nested capture groups in named groups #4641
RyanL1997 merged 7 commits into
opensearch-project:mainfrom
RyanL1997:rex-extract-fix

Conversation

@RyanL1997

@RyanL1997RyanL1997 commented Oct 22, 2025

Copy link
Copy Markdown
Collaborator

Description

Fix unexpected shift of extraction for rex with nested capture groups in named groups

The rex command in PPL had a critical bug when using named capture groups that contained nested unnamed groups. This caused extracted field values to shift by one position, producing incorrect results.

  • Root Cause: Code used sequential indices 1, 2, 3... but nested groups create non-sequential indices 1, 3, 5...
  • Solution: Bypass index calculation entirely by using Java's native named group extraction (matcher.group(groupName))

Example of the Bug

Query:

curl -X POST "localhost:9200/_plugins/_ppl" \
-H "Content-Type: application/json" \
-d '{ "query": "source=accounts | rex field=email \"(?<user>(amber|hattie|nanette)[a-z]*)@(?<domain>(pyrami|netagy|quility))\\.(?<tld>(com|org))\" | fields user, domain, tld | head 1" }'

Expected Result (correct):

["amberduke", "pyrami", "com"]

Actual Result (wrong):

["amberduke", "amber", "pyrami"]

Root Cause

When Java's regex engine processes the pattern (?<user>(amber|hattie))[a-z]*, it assigns group numbers to ALL capture groups (named and unnamed):

Pattern: (?<user>(amber|hattie))[a-z]*@(?<domain>(pyrami|netagy))\.(?<tld>(com|org))

Group Assignment:

  • Group 0: Entire match
  • Group 1: (?<user>...) ← Named group "user"
  • Group 2: (amber|hattie) ← Unnamed nested group
  • Group 3: (?<domain>...) ← Named group "domain"
  • Group 4: (pyrami|netagy) ← Unnamed nested group
  • Group 5: (?<tld>...) ← Named group "tld"
  • Group 6: (com|org) ← Unnamed nested group

Before the fix, the bug is in CalciteRelNodeVisitor.java (lines 265-321). The code does:

List<String> namedGroups = RegexCommonUtils.getNamedGroupCandidates(patternStr);
// namedGroups = ["user", "domain", "tld"]for (inti = 0; i < namedGroups.size(); i++) {
extractCall = PPLFuncImpTable.INSTANCE.resolve(
context.rexBuilder,
BuiltinFunctionName.REX_EXTRACT,
fieldRex,
context.rexBuilder.makeLiteral(patternStr),
context.relBuilder.literal(i + 1)); // ← WRONG: Assumes sequential named groups// ...
}

The code assumes named groups are at indices 1, 2, 3, ... but the actual indices are 1, 3, 5, ... due to the unnamed nested groups.

With the above buggy logic:

  • REX_EXTRACT(field, pattern, 1) → Gets Group 1 (?<user>...) = "amberduke" → CORRECT
  • REX_EXTRACT(field, pattern, 2) → Gets Group 2 (amber|hattie) = "amber" → WRONG
  • REX_EXTRACT(field, pattern, 3) → Gets Group 3 (?<domain>...) = "pyrami" → WRONG

The second and third extractions are off by one group because they hit the unnamed nested groups.

LogicalProject(
user=[REX_EXTRACT($7, '(?<user>(amber|hattie))[a-z]*@(?<domain>(pyrami|netagy))\.(?<tld>(com|org))', 1)],
domain=[REX_EXTRACT($7, '(?<user>(amber|hattie))[a-z]*@(?<domain>(pyrami|netagy))\.(?<tld>(com|org))', 2)], -- Wrong!
tld=[REX_EXTRACT($7, '(?<user>(amber|hattie))[a-z]*@(?<domain>(pyrami|netagy))\.(?<tld>(com|org))', 3)] -- Wrong!
)

Related Issues

Check List

  • New functionality includes testing.
  • New functionality has been documented.
  • New functionality has javadoc added.
  • New functionality has a user manual doc added.
  • New PPL command checklist all confirmed.
  • API changes companion pull request created.
  • Commits are signed per the DCO using --signoff or -s.
  • Public documentation issue/PR created.

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.
For more information on following Developer Certificate of Origin and signing off your commits, please check here.

matchCount++;
} else {
// If extractor returns null, it might indicate an error (like invalid group name)
// Stop processing to avoid infinite loop

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm currently thinking about adding an error handling here

Signed-off-by: Jialiang Liang <jiallian@amazon.com>
Signed-off-by: Jialiang Liang <jiallian@amazon.com>
Signed-off-by: Jialiang Liang <jiallian@amazon.com>
Signed-off-by: Jialiang Liang <jiallian@amazon.com>
Signed-off-by: Jialiang Liang <jiallian@amazon.com>
Signed-off-by: Jialiang Liang <jiallian@amazon.com>
@RyanL1997RyanL1997 changed the title [BugFix] Fix the off-by-one error for rex with nested capture groups in named groups [BugFix] Fix unexpected shift of extraction for rex with nested capture groups in named groups Oct 25, 2025
Swiddis
Swiddis previously approved these changes Oct 28, 2025
fieldRex,
context.rexBuilder.makeLiteral(patternStr),
context.relBuilder.literal(i + 1),
context.rexBuilder.makeLiteral(groupName),

@dai-chendai-chenOct 28, 2025

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I found there is a namedGroups() API in Matcher (since JDK 20?). If we can get correct index here, we don't need to modify the UDFs below? Alternatively we can move capture name -> index logic here from UDFs?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That is correct - the core issue is matching named groups to their correct indices, and Pattern.namedGroups() would be the perfect solution. However, I discovered that we're blocked by a
compatibility constraint:

  • Pattern.namedGroups() was introduced in JDK 20
  • We need backward compatibility with JDK 11/17 - for 2.19-dev

I agree that directly leveraging the Pattern.namedGroups() is the right architectural approach - we should definitely migrate to it when we fully upgrade to JDK 20+. At that point, it would be a simple one-line change in CalciteRelNodeVisitor.

Signed-off-by: Jialiang Liang <jiallian@amazon.com>
"Rex pattern must contain at least one named capture group");
}

// TODO: Once JDK 20+ is supported, consider using Pattern.namedGroups() API for more efficient

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

as for now, I added a TODO here @dai-chen

@RyanL1997
RyanL1997 merged commit 0c1ec27 into opensearch-project:mainOct 29, 2025
51 of 52 checks passed
@RyanL1997
RyanL1997 deleted the rex-extract-fix branch October 29, 2025 00:34
opensearch-trigger-botBot pushed a commit that referenced this pull request Oct 29, 2025
…ture groups in named groups (#4641)
(cherry picked from commit 0c1ec27)
Signed-off-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
RyanL1997 pushed a commit that referenced this pull request Oct 29, 2025
…ture groups in named groups (#4641) (#4692)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
expani pushed a commit to vinaykpud/sql that referenced this pull request Nov 4, 2025
sandeshkr419 added a commit to sandeshkr419/sql that referenced this pull request Dec 3, 2025
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Simeon Widdis <sawiddis@amazon.com>
Co-authored-by: Manasvini B S <manasvis@amazon.com>
Co-authored-by: opensearch-ci-bot <opensearch-infra@amazon.com>
Co-authored-by: Louis Chu <clingzhi@amazon.com>
Co-authored-by: Chen Dai <daichen@amazon.com>
Co-authored-by: Mebsina <cnoramut@gmail.com>
Co-authored-by: Yuanchun Shen <yuanchu@amazon.com>
Co-authored-by: opensearch-trigger-bot[bot] <98922864+opensearch-trigger-bot[bot]@users.noreply.github.com>
Co-authored-by: Kai Huang <105710027+ahkcs@users.noreply.github.com>
Co-authored-by: Peng Huo <penghuo@gmail.com>
Co-authored-by: Alexey Temnikov <alexey.temnikov@improving.com>
Co-authored-by: Riley Jerger <214163063+RileyJergerAmazon@users.noreply.github.com>
Co-authored-by: Tomoyuki MORITA <moritato@amazon.com>
Co-authored-by: Lantao Jin <ltjin@amazon.com>
Co-authored-by: Songkan Tang <songkant@amazon.com>
Co-authored-by: qianheng <qianheng@amazon.com>
Co-authored-by: Simeon Widdis <sawiddis@gmail.com>
Co-authored-by: Xinyuan Lu <xinyual@amazon.com>
Co-authored-by: Jialiang Liang <jiallian@amazon.com>
Co-authored-by: Peter Zhu <zhujiaxi@amazon.com>
Co-authored-by: Vinay Krishna Pudyodu <vinkrish.neo@gmail.com>
Co-authored-by: expani <anijainc@amazon.com>
Co-authored-by: expani1729 <110471048+expani@users.noreply.github.com>
Co-authored-by: Vamsi Manohar <reddyvam@amazon.com>
Co-authored-by: ritvibhatt <53196324+ritvibhatt@users.noreply.github.com>
Co-authored-by: Xinyu Hao <75524174+ishaoxy@users.noreply.github.com>
Co-authored-by: Marc Handalian <marc.handalian@gmail.com>
Co-authored-by: Marc Handalian <handalm@amazon.com>
Fix join type ambiguous issue when specify the join type with sql-like join criteria (opensearch-project#4474)
Fix issue 4441 (opensearch-project#4449)
Fix missing keywordsCanBeId (opensearch-project#4491)
Fix the bug of explicit makeNullLiteral for UDT fields (opensearch-project#4475)
Fix mapping after aggregation push down (opensearch-project#4500)
Fix percentile bug (opensearch-project#4539)
Fix JsonExtractAllFunctionIT failure (opensearch-project#4556)
Fix sort push down into agg after project already pushed (opensearch-project#4546)
Fix push down failure for min/max on derived field (opensearch-project#4572)
Fix compile issue in main (opensearch-project#4608)
Fix filter parsing failure on date fields with non-default format (opensearch-project#4616)
Fix bin nested fields issue (opensearch-project#4606)
Fix: Support Alias Fields in MIN, MAX, FIRST, LAST, and TAKE Aggregations (opensearch-project#4621)
fix rename issue (opensearch-project#4670)
Fixes for `Multisearch` and `Append` command (opensearch-project#4512)
Fix asc/desc keyword behavior for sort command (opensearch-project#4651)
Fix] Fix unexpected shift of extraction for `rex` with nested capture groups in named groups (opensearch-project#4641)
Fix CVE-2025-48924 (opensearch-project#4665)
Fix sub-fields accessing of generated structs (opensearch-project#4683)
Fix] Incorrect Field Index Mapping in AVG to SUM/COUNT Conversion (opensearch-project#15)
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

backport 2.19-devbugSomething isn't workingbugFixPPLPiped processing language

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[BUG] rex command off-by-one error with nested capture groups in named groups

3 participants

@RyanL1997@Swiddis@dai-chen
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

[BugFix] Fix unexpected shift of extraction for rex with nested capture groups in named groups - #4641

Merged
RyanL1997 merged 7 commits into
opensearch-project:mainfrom
RyanL1997:rex-extract-fix
Oct 29, 2025
Merged

[BugFix] Fix unexpected shift of extraction for rex with nested capture groups in named groups #4641
RyanL1997 merged 7 commits into
opensearch-project:mainfrom
RyanL1997:rex-extract-fix

Conversation

@RyanL1997

@RyanL1997RyanL1997 commented Oct 22, 2025

Copy link
Copy Markdown
Collaborator

Description

Fix unexpected shift of extraction for rex with nested capture groups in named groups

The rex command in PPL had a critical bug when using named capture groups that contained nested unnamed groups. This caused extracted field values to shift by one position, producing incorrect results.

  • Root Cause: Code used sequential indices 1, 2, 3... but nested groups create non-sequential indices 1, 3, 5...
  • Solution: Bypass index calculation entirely by using Java's native named group extraction (matcher.group(groupName))

Example of the Bug

Query:

curl -X POST "localhost:9200/_plugins/_ppl" \
-H "Content-Type: application/json" \
-d '{ "query": "source=accounts | rex field=email \"(?<user>(amber|hattie|nanette)[a-z]*)@(?<domain>(pyrami|netagy|quility))\\.(?<tld>(com|org))\" | fields user, domain, tld | head 1" }'

Expected Result (correct):

["amberduke", "pyrami", "com"]

Actual Result (wrong):

["amberduke", "amber", "pyrami"]

Root Cause

When Java's regex engine processes the pattern (?<user>(amber|hattie))[a-z]*, it assigns group numbers to ALL capture groups (named and unnamed):

Pattern: (?<user>(amber|hattie))[a-z]*@(?<domain>(pyrami|netagy))\.(?<tld>(com|org))

Group Assignment:

  • Group 0: Entire match
  • Group 1: (?<user>...) ← Named group "user"
  • Group 2: (amber|hattie) ← Unnamed nested group
  • Group 3: (?<domain>...) ← Named group "domain"
  • Group 4: (pyrami|netagy) ← Unnamed nested group
  • Group 5: (?<tld>...) ← Named group "tld"
  • Group 6: (com|org) ← Unnamed nested group

Before the fix, the bug is in CalciteRelNodeVisitor.java (lines 265-321). The code does:

List<String> namedGroups = RegexCommonUtils.getNamedGroupCandidates(patternStr);
// namedGroups = ["user", "domain", "tld"]for (inti = 0; i < namedGroups.size(); i++) {
extractCall = PPLFuncImpTable.INSTANCE.resolve(
context.rexBuilder,
BuiltinFunctionName.REX_EXTRACT,
fieldRex,
context.rexBuilder.makeLiteral(patternStr),
context.relBuilder.literal(i + 1)); // ← WRONG: Assumes sequential named groups// ...
}

The code assumes named groups are at indices 1, 2, 3, ... but the actual indices are 1, 3, 5, ... due to the unnamed nested groups.

With the above buggy logic:

  • REX_EXTRACT(field, pattern, 1) → Gets Group 1 (?<user>...) = "amberduke" → CORRECT
  • REX_EXTRACT(field, pattern, 2) → Gets Group 2 (amber|hattie) = "amber" → WRONG
  • REX_EXTRACT(field, pattern, 3) → Gets Group 3 (?<domain>...) = "pyrami" → WRONG

The second and third extractions are off by one group because they hit the unnamed nested groups.

LogicalProject(
user=[REX_EXTRACT($7, '(?<user>(amber|hattie))[a-z]*@(?<domain>(pyrami|netagy))\.(?<tld>(com|org))', 1)],
domain=[REX_EXTRACT($7, '(?<user>(amber|hattie))[a-z]*@(?<domain>(pyrami|netagy))\.(?<tld>(com|org))', 2)], -- Wrong!
tld=[REX_EXTRACT($7, '(?<user>(amber|hattie))[a-z]*@(?<domain>(pyrami|netagy))\.(?<tld>(com|org))', 3)] -- Wrong!
)

Related Issues

Check List

  • New functionality includes testing.
  • New functionality has been documented.
  • New functionality has javadoc added.
  • New functionality has a user manual doc added.
  • New PPL command checklist all confirmed.
  • API changes companion pull request created.
  • Commits are signed per the DCO using --signoff or -s.
  • Public documentation issue/PR created.

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.
For more information on following Developer Certificate of Origin and signing off your commits, please check here.

matchCount++;
} else {
// If extractor returns null, it might indicate an error (like invalid group name)
// Stop processing to avoid infinite loop

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm currently thinking about adding an error handling here

Signed-off-by: Jialiang Liang <jiallian@amazon.com>
Signed-off-by: Jialiang Liang <jiallian@amazon.com>
Signed-off-by: Jialiang Liang <jiallian@amazon.com>
Signed-off-by: Jialiang Liang <jiallian@amazon.com>
Signed-off-by: Jialiang Liang <jiallian@amazon.com>
Signed-off-by: Jialiang Liang <jiallian@amazon.com>
@RyanL1997RyanL1997 changed the title [BugFix] Fix the off-by-one error for rex with nested capture groups in named groups [BugFix] Fix unexpected shift of extraction for rex with nested capture groups in named groups Oct 25, 2025
Swiddis
Swiddis previously approved these changes Oct 28, 2025
fieldRex,
context.rexBuilder.makeLiteral(patternStr),
context.relBuilder.literal(i + 1),
context.rexBuilder.makeLiteral(groupName),

@dai-chendai-chenOct 28, 2025

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I found there is a namedGroups() API in Matcher (since JDK 20?). If we can get correct index here, we don't need to modify the UDFs below? Alternatively we can move capture name -> index logic here from UDFs?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That is correct - the core issue is matching named groups to their correct indices, and Pattern.namedGroups() would be the perfect solution. However, I discovered that we're blocked by a
compatibility constraint:

  • Pattern.namedGroups() was introduced in JDK 20
  • We need backward compatibility with JDK 11/17 - for 2.19-dev

I agree that directly leveraging the Pattern.namedGroups() is the right architectural approach - we should definitely migrate to it when we fully upgrade to JDK 20+. At that point, it would be a simple one-line change in CalciteRelNodeVisitor.

Signed-off-by: Jialiang Liang <jiallian@amazon.com>
"Rex pattern must contain at least one named capture group");
}

// TODO: Once JDK 20+ is supported, consider using Pattern.namedGroups() API for more efficient

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

as for now, I added a TODO here @dai-chen

@RyanL1997
RyanL1997 merged commit 0c1ec27 into opensearch-project:mainOct 29, 2025
51 of 52 checks passed
@RyanL1997
RyanL1997 deleted the rex-extract-fix branch October 29, 2025 00:34
opensearch-trigger-botBot pushed a commit that referenced this pull request Oct 29, 2025
…ture groups in named groups (#4641)
(cherry picked from commit 0c1ec27)
Signed-off-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
RyanL1997 pushed a commit that referenced this pull request Oct 29, 2025
…ture groups in named groups (#4641) (#4692)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
expani pushed a commit to vinaykpud/sql that referenced this pull request Nov 4, 2025
sandeshkr419 added a commit to sandeshkr419/sql that referenced this pull request Dec 3, 2025
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Simeon Widdis <sawiddis@amazon.com>
Co-authored-by: Manasvini B S <manasvis@amazon.com>
Co-authored-by: opensearch-ci-bot <opensearch-infra@amazon.com>
Co-authored-by: Louis Chu <clingzhi@amazon.com>
Co-authored-by: Chen Dai <daichen@amazon.com>
Co-authored-by: Mebsina <cnoramut@gmail.com>
Co-authored-by: Yuanchun Shen <yuanchu@amazon.com>
Co-authored-by: opensearch-trigger-bot[bot] <98922864+opensearch-trigger-bot[bot]@users.noreply.github.com>
Co-authored-by: Kai Huang <105710027+ahkcs@users.noreply.github.com>
Co-authored-by: Peng Huo <penghuo@gmail.com>
Co-authored-by: Alexey Temnikov <alexey.temnikov@improving.com>
Co-authored-by: Riley Jerger <214163063+RileyJergerAmazon@users.noreply.github.com>
Co-authored-by: Tomoyuki MORITA <moritato@amazon.com>
Co-authored-by: Lantao Jin <ltjin@amazon.com>
Co-authored-by: Songkan Tang <songkant@amazon.com>
Co-authored-by: qianheng <qianheng@amazon.com>
Co-authored-by: Simeon Widdis <sawiddis@gmail.com>
Co-authored-by: Xinyuan Lu <xinyual@amazon.com>
Co-authored-by: Jialiang Liang <jiallian@amazon.com>
Co-authored-by: Peter Zhu <zhujiaxi@amazon.com>
Co-authored-by: Vinay Krishna Pudyodu <vinkrish.neo@gmail.com>
Co-authored-by: expani <anijainc@amazon.com>
Co-authored-by: expani1729 <110471048+expani@users.noreply.github.com>
Co-authored-by: Vamsi Manohar <reddyvam@amazon.com>
Co-authored-by: ritvibhatt <53196324+ritvibhatt@users.noreply.github.com>
Co-authored-by: Xinyu Hao <75524174+ishaoxy@users.noreply.github.com>
Co-authored-by: Marc Handalian <marc.handalian@gmail.com>
Co-authored-by: Marc Handalian <handalm@amazon.com>
Fix join type ambiguous issue when specify the join type with sql-like join criteria (opensearch-project#4474)
Fix issue 4441 (opensearch-project#4449)
Fix missing keywordsCanBeId (opensearch-project#4491)
Fix the bug of explicit makeNullLiteral for UDT fields (opensearch-project#4475)
Fix mapping after aggregation push down (opensearch-project#4500)
Fix percentile bug (opensearch-project#4539)
Fix JsonExtractAllFunctionIT failure (opensearch-project#4556)
Fix sort push down into agg after project already pushed (opensearch-project#4546)
Fix push down failure for min/max on derived field (opensearch-project#4572)
Fix compile issue in main (opensearch-project#4608)
Fix filter parsing failure on date fields with non-default format (opensearch-project#4616)
Fix bin nested fields issue (opensearch-project#4606)
Fix: Support Alias Fields in MIN, MAX, FIRST, LAST, and TAKE Aggregations (opensearch-project#4621)
fix rename issue (opensearch-project#4670)
Fixes for `Multisearch` and `Append` command (opensearch-project#4512)
Fix asc/desc keyword behavior for sort command (opensearch-project#4651)
Fix] Fix unexpected shift of extraction for `rex` with nested capture groups in named groups (opensearch-project#4641)
Fix CVE-2025-48924 (opensearch-project#4665)
Fix sub-fields accessing of generated structs (opensearch-project#4683)
Fix] Incorrect Field Index Mapping in AVG to SUM/COUNT Conversion (opensearch-project#15)
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

backport 2.19-devbugSomething isn't workingbugFixPPLPiped processing language

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[BUG] rex command off-by-one error with nested capture groups in named groups

3 participants

@RyanL1997@Swiddis@dai-chen