feat(sql): Support RANK/DENSE_RANK for unified SQL - #5720

Merged
dai-chen merged 1 commit into
opensearch-project:mainfrom
dai-chen:support-rank-dense-rank-calcite
Aug 28, 2026
Merged

feat(sql): Support RANK/DENSE_RANK for unified SQL#5720
dai-chen merged 1 commit into
opensearch-project:mainfrom
dai-chen:support-rank-dense-rank-calcite

Conversation

@dai-chen

@dai-chendai-chen commented Aug 24, 2026

Copy link
Copy Markdown
Collaborator

Description

RANK() and DENSE_RANK() were already accepted by the SQL grammar and declared as built-in function names, but were never registered as window functions, so planning failed with Unexpected window function: RANK. This PR registers both and maps them to Calcite's standard RANK and DENSE_RANK operators. No grammar, lexer or AST changes were needed.

Related Issues

Part of #5248

Check List

  • New functionality includes testing.
  • New functionality has been documented.
  • New functionality has javadoc added.
  • New functionality has a user manual doc added.
  • New PPL command checklist all confirmed.
  • API changes companion pull request created.
  • Commits are signed per the DCO using --signoff or -s.
  • Public documentation issue/PR created.

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.
For more information on following Developer Certificate of Origin and signing off your commits, please check here.

@dai-chendai-chen self-assigned this Aug 24, 2026
@github-actions

github-actionsBot commented Aug 24, 2026

Copy link
Copy Markdown
Contributor

PR Reviewer Guide 🔍

(Review updated until commit 544e467)

Here are some key observations to aid the review process:

🧪 PR contains tests
🔒 No security concerns identified
✅ No TODO sections
🔀 No multiple PR themes
⚡ No major issues detected

@github-actions

github-actionsBot commented Aug 24, 2026

Copy link
Copy Markdown
Contributor

PR Code Suggestions ✨

Latest suggestions up to 544e467

Explore these optional code suggestions:

CategorySuggestion Impact
General
Explicitly pass null for frame bounds

Since RANK and DENSE_RANK disallow framing, passing lowerBound and upperBound
parameters may cause unexpected behavior or errors. Consider passing null for both
bounds to explicitly indicate no framing is used.

core/src/main/java/org/opensearch/sql/calcite/utils/PlanUtils.java [236-251]

 case RANK:
return withOver(
context.relBuilder.aggregateCall(SqlStdOperatorTable.RANK),
partitions,
orderKeys,
false,
- lowerBound,- upperBound);+ null,+ null);
case DENSE_RANK:
return withOver(
context.relBuilder.aggregateCall(SqlStdOperatorTable.DENSE_RANK),
partitions,
orderKeys,
false,
- lowerBound,- upperBound);+ null,+ null);
Suggestion importance[1-10]: 7

__

Why: The suggestion correctly identifies that RANK and DENSE_RANK disallow framing, and the comment in the code confirms this. Passing null instead of lowerBound and upperBound would make the intent more explicit and prevent potential issues, though the current implementation may already handle this correctly through normalization.

Medium

Previous suggestions

Suggestions up to commit 77ae6e0
CategorySuggestion Impact
Possible issue
Explicitly pass null for frame bounds

Since RANK and DENSE_RANK disallow framing, passing lowerBound and upperBound
parameters may cause unexpected behavior or errors. Consider explicitly passing null
for both bounds to ensure proper handling of these ranking functions.

core/src/main/java/org/opensearch/sql/calcite/utils/PlanUtils.java [236-251]

 case RANK:
return withOver(
context.relBuilder.aggregateCall(SqlStdOperatorTable.RANK),
partitions,
orderKeys,
false,
- lowerBound,- upperBound);+ null,+ null);
case DENSE_RANK:
return withOver(
context.relBuilder.aggregateCall(SqlStdOperatorTable.DENSE_RANK),
partitions,
orderKeys,
false,
- lowerBound,- upperBound);+ null,+ null);
Suggestion importance[1-10]: 4

__

Why: While the suggestion correctly identifies that RANK and DENSE_RANK disallow framing, the comment in the PR already acknowledges this ("Calcite rank operators disallow framing, so the ROWS/RANGE flag below is normalized away"). The current implementation passes lowerBound and upperBound which are likely normalized by Calcite. Explicitly passing null would be slightly clearer but is not critical since the normalization handles this.

Low
Suggestions up to commit 747d1b6
CategorySuggestion Impact
Possible issue
Validate ORDER BY clause presence

The RANK and DENSE_RANK window functions require an ORDER BY clause to function
correctly. Consider validating that orderKeys is not empty before creating the
aggregate call to prevent runtime errors or unexpected behavior when no ordering is
specified.

core/src/main/java/org/opensearch/sql/calcite/utils/PlanUtils.java [236-251]

 case RANK:
+ if (orderKeys.isEmpty()) {+ throw new IllegalArgumentException("RANK requires ORDER BY clause");+ }
return withOver(
context.relBuilder.aggregateCall(SqlStdOperatorTable.RANK),
partitions,
orderKeys,
false,
lowerBound,
upperBound);
case DENSE_RANK:
+ if (orderKeys.isEmpty()) {+ throw new IllegalArgumentException("DENSE_RANK requires ORDER BY clause");+ }
return withOver(
context.relBuilder.aggregateCall(SqlStdOperatorTable.DENSE_RANK),
partitions,
orderKeys,
false,
lowerBound,
upperBound);
Suggestion importance[1-10]: 5

__

Why: While RANK and DENSE_RANK do require an ORDER BY clause semantically, this validation may already be handled at the SQL parsing/validation layer. Adding redundant validation here could be useful for defensive programming, but without evidence of missing validation upstream, this is a moderate improvement for robustness rather than fixing a critical bug.

Low
Suggestions up to commit 352ffbd
CategorySuggestion Impact
General
Use Set for multiple condition checks

Consider using a Set or EnumSet for checking multiple function names instead of
chained OR conditions. This improves readability and makes it easier to add more
functions in the future.

core/src/main/java/org/opensearch/sql/calcite/CalciteRexNodeVisitor.java [775-777]

-if (functionName == BuiltinFunctionName.ROW_NUMBER- || functionName == BuiltinFunctionName.RANK- || functionName == BuiltinFunctionName.DENSE_RANK) {+private static final Set<BuiltinFunctionName> NO_FIELD_WINDOW_FUNCTIONS = + EnumSet.of(BuiltinFunctionName.ROW_NUMBER, BuiltinFunctionName.RANK, BuiltinFunctionName.DENSE_RANK);+if (NO_FIELD_WINDOW_FUNCTIONS.contains(functionName)) {+
Suggestion importance[1-10]: 4

__

Why: While using an EnumSet would improve maintainability and readability, the current chained OR condition with three items is still acceptable and clear. The suggestion is valid but offers only a moderate improvement in code style rather than fixing a critical issue.

Low

@dai-chen
dai-chenforce-pushed the support-rank-dense-rank-calcite branch from 352ffbd to 747d1b6CompareAugust 25, 2026 00:21
@github-actions

Copy link
Copy Markdown
Contributor

Persistent review updated to latest commit 747d1b6

@dai-chen
dai-chenforce-pushed the support-rank-dense-rank-calcite branch from 747d1b6 to 77ae6e0CompareAugust 25, 2026 18:17
@github-actions

Copy link
Copy Markdown
Contributor

Persistent review updated to latest commit 77ae6e0

Both were accepted by the SQL grammar and declared as
BuiltinFunctionName constants, but missing from WINDOW_FUNC_MAPPING, so
visitWindowFunction rejected them with "Unexpected window function".
Register both, extend the ROW_NUMBER bypass of aggregate signature
validation since they likewise take no arguments, and lower them to
SqlStdOperatorTable.RANK and DENSE_RANK.
WINDOW_FUNC_MAPPING is shared, so PPL eventstats/streamstats now resolve
these functions as well.
Related to opensearch-project#5168
Signed-off-by: Chen Dai <daichen@amazon.com>
@dai-chen
dai-chenforce-pushed the support-rank-dense-rank-calcite branch from 77ae6e0 to 544e467CompareAugust 26, 2026 22:13
@github-actions

Copy link
Copy Markdown
Contributor

Persistent review updated to latest commit 544e467

@dai-chendai-chen changed the title feat(sql): Support RANK/DENSE_RANK in unified SQLfeat(sql): Support RANK/DENSE_RANK for unified SQLAug 26, 2026
@dai-chen
dai-chen marked this pull request as ready for review August 26, 2026 23:10
gingeekrishna added a commit to gingeekrishna/sql that referenced this pull request Aug 28, 2026
WINDOW_FUNC_MAPPING (used by eventstats/streamstats) never supported
rank/dense_rank, but PPL's scalarWindowFunctionName grammar rule still
accepted the tokens, so `eventstats rank()` reached
CalciteRexNodeVisitor#visitWindowFunction and failed there with the
"not supported" message this PR improves. Meanwhile SQL's grammar
already accepts RANK()/DENSE_RANK() OVER (...), and opensearch-project#5720 is adding
real support for them on the SQL side via the same shared visitor.
Remove RANK/DENSE_RANK from scalarWindowFunctionName so PPL rejects
them at parse time instead of falling through to the shared
SQL/PPL visitor - this keeps the language separation at the parser
rather than relying on a WINDOW_FUNC_MAPPING check in shared planner
code, and avoids PPL silently gaining rank/dense_rank as a side effect
of opensearch-project#5720 registering them for SQL.
Updates the eventstats/streamstats tests added earlier in this PR to
expect a SyntaxCheckException (parse-time) instead of the semantic
"not supported" error, and switches the unrelated
visitWindowFunction-rejection unit test from rank() to percent_rank(),
which remains unsupported and still exercises that code path.
Signed-off-by: Radhakrishnan P <gingeekrishna@gmail.com>

@RyanL1997RyanL1997 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM. Just would like to double check do we need to add doc test/doc for this?

@RyanL1997

Copy link
Copy Markdown
Collaborator

In addtion, I found out that

The following isn't something you introduced in this PR. The existing ROW_NUMBER case has the same shape, and ROW_NUMBER() OVER (ORDER BY age) already returns 1 for every row. The new RANK/DENSE_RANK cases inherit it.

Behaviour:

SELECT age, RANK() OVER (ORDER BY age) FROM employees
actual → 4, 4, 4, 4 (the partition size)
expected → 1, 2, 3, 4

Why: the new cases forward the caller's lowerBound/upperBound unchanged, and for SQL those default to WindowFrame.rowsUnbounded() = UNBOUNDED PRECEDING … UNBOUNDED FOLLOWING, so every row is ranked over the whole partition. Calcite's own SqlToRelConverter.convertOver forces UNBOUNDED PRECEDING … CURRENT ROW for operators with allowsFraming() == false; this branch needs to do the same.

One knock-on: the plan assertions can't catch it. allowsFraming() only suppresses the frame in the printed digest, not in the executed Window.Group — so RANK() OVER (ORDER BY $2 NULLS FIRST) prints identically whichever frame was built. Confirming this needs an execution-level assertion.

@dai-chen
dai-chen merged commit c4f27b2 into opensearch-project:mainAug 28, 2026
42 checks passed
gingeekrishna added a commit to gingeekrishna/sql that referenced this pull request Aug 30, 2026
WINDOW_FUNC_MAPPING (used by eventstats/streamstats) never supported
rank/dense_rank, but PPL's scalarWindowFunctionName grammar rule still
accepted the tokens, so `eventstats rank()` reached
CalciteRexNodeVisitor#visitWindowFunction and failed there with the
"not supported" message this PR improves. Meanwhile SQL's grammar
already accepts RANK()/DENSE_RANK() OVER (...), and opensearch-project#5720 is adding
real support for them on the SQL side via the same shared visitor.
Remove RANK/DENSE_RANK from scalarWindowFunctionName so PPL rejects
them at parse time instead of falling through to the shared
SQL/PPL visitor - this keeps the language separation at the parser
rather than relying on a WINDOW_FUNC_MAPPING check in shared planner
code, and avoids PPL silently gaining rank/dense_rank as a side effect
of opensearch-project#5720 registering them for SQL.
Updates the eventstats/streamstats tests added earlier in this PR to
expect a SyntaxCheckException (parse-time) instead of the semantic
"not supported" error, and switches the unrelated
visitWindowFunction-rejection unit test from rank() to percent_rank(),
which remains unsupported and still exercises that code path.
Signed-off-by: Radhakrishnan P <gingeekrishna@gmail.com>
gingeekrishna added a commit to gingeekrishna/sql that referenced this pull request Sep 2, 2026
…layer
opensearch-project#5720 added rank/dense_rank to the WINDOW_FUNC_MAPPING shared by both
SQL's RANK()/DENSE_RANK() OVER (...) and PPL's eventstats/streamstats,
which silently enabled them for eventstats/streamstats too:
testRankingWindowFunctionsUnsupportedInEventstats/InStreamstats
(added earlier in this PR) started failing with "expected
ResponseException to be thrown, but nothing was thrown" once opensearch-project#5720
merged, since eventstats rank() now builds a real window call instead
of hitting the "not supported" check.
That's a real gap, not just a test artifact: PPL's eventstats/
streamstats grammar has no ORDER BY syntax at all, so ranking has no
defined ordering to rank by there - unlike SQL's OVER(), which at
least has (optional) ORDER BY in its own clause.
A prior commit on this branch tried fixing this by removing RANK/
DENSE_RANK from the PPL grammar entirely, rejecting them at parse
time. That broke a different, cross-repo contract: OpenSearch-
Dashboards' PPL linter (validated by the "PPL grammar compatibility"
CI check) expects `eventstats rank()` to parse successfully and be
flagged by a semantic-layer diagnostic instead, so it could no longer
produce that diagnostic once the query stopped parsing. That commit
was reverted.
Fix this at the semantic layer instead, where it belongs: PPL only
ever reaches CalciteRexNodeVisitor#visitWindowFunction through
eventstats/streamstats (no other PPL syntax builds a WindowFunction
node), so context.queryType == PPL is an exact, unambiguous signal for
"this is an eventstats/streamstats call". Filter rank/dense_rank out
of the WINDOW_FUNC_MAPPING lookup specifically when queryType is PPL,
so they fall through to the existing "not supported in eventstats/
streamstats" error - exactly the pre-opensearch-project#5720 behavior - while leaving
SQL's RANK()/DENSE_RANK() OVER (...) handling (and everything else)
completely untouched.
Signed-off-by: Radhakrishnan P <gingeekrishna@gmail.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

analytic-engineenhancementNew feature or requestSQL

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@dai-chen@RyanL1997
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

feat(sql): Support RANK/DENSE_RANK for unified SQL - #5720

Merged
dai-chen merged 1 commit into
opensearch-project:mainfrom
dai-chen:support-rank-dense-rank-calcite
Aug 28, 2026
Merged

feat(sql): Support RANK/DENSE_RANK for unified SQL#5720
dai-chen merged 1 commit into
opensearch-project:mainfrom
dai-chen:support-rank-dense-rank-calcite

Conversation

@dai-chen

@dai-chendai-chen commented Aug 24, 2026

Copy link
Copy Markdown
Collaborator

Description

RANK() and DENSE_RANK() were already accepted by the SQL grammar and declared as built-in function names, but were never registered as window functions, so planning failed with Unexpected window function: RANK. This PR registers both and maps them to Calcite's standard RANK and DENSE_RANK operators. No grammar, lexer or AST changes were needed.

Related Issues

Part of #5248

Check List

  • New functionality includes testing.
  • New functionality has been documented.
  • New functionality has javadoc added.
  • New functionality has a user manual doc added.
  • New PPL command checklist all confirmed.
  • API changes companion pull request created.
  • Commits are signed per the DCO using --signoff or -s.
  • Public documentation issue/PR created.

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.
For more information on following Developer Certificate of Origin and signing off your commits, please check here.

@dai-chendai-chen self-assigned this Aug 24, 2026
@github-actions

github-actionsBot commented Aug 24, 2026

Copy link
Copy Markdown
Contributor

PR Reviewer Guide 🔍

(Review updated until commit 544e467)

Here are some key observations to aid the review process:

🧪 PR contains tests
🔒 No security concerns identified
✅ No TODO sections
🔀 No multiple PR themes
⚡ No major issues detected

@github-actions

github-actionsBot commented Aug 24, 2026

Copy link
Copy Markdown
Contributor

PR Code Suggestions ✨

Latest suggestions up to 544e467

Explore these optional code suggestions:

CategorySuggestion Impact
General
Explicitly pass null for frame bounds

Since RANK and DENSE_RANK disallow framing, passing lowerBound and upperBound
parameters may cause unexpected behavior or errors. Consider passing null for both
bounds to explicitly indicate no framing is used.

core/src/main/java/org/opensearch/sql/calcite/utils/PlanUtils.java [236-251]

 case RANK:
return withOver(
context.relBuilder.aggregateCall(SqlStdOperatorTable.RANK),
partitions,
orderKeys,
false,
- lowerBound,- upperBound);+ null,+ null);
case DENSE_RANK:
return withOver(
context.relBuilder.aggregateCall(SqlStdOperatorTable.DENSE_RANK),
partitions,
orderKeys,
false,
- lowerBound,- upperBound);+ null,+ null);
Suggestion importance[1-10]: 7

__

Why: The suggestion correctly identifies that RANK and DENSE_RANK disallow framing, and the comment in the code confirms this. Passing null instead of lowerBound and upperBound would make the intent more explicit and prevent potential issues, though the current implementation may already handle this correctly through normalization.

Medium

Previous suggestions

Suggestions up to commit 77ae6e0
CategorySuggestion Impact
Possible issue
Explicitly pass null for frame bounds

Since RANK and DENSE_RANK disallow framing, passing lowerBound and upperBound
parameters may cause unexpected behavior or errors. Consider explicitly passing null
for both bounds to ensure proper handling of these ranking functions.

core/src/main/java/org/opensearch/sql/calcite/utils/PlanUtils.java [236-251]

 case RANK:
return withOver(
context.relBuilder.aggregateCall(SqlStdOperatorTable.RANK),
partitions,
orderKeys,
false,
- lowerBound,- upperBound);+ null,+ null);
case DENSE_RANK:
return withOver(
context.relBuilder.aggregateCall(SqlStdOperatorTable.DENSE_RANK),
partitions,
orderKeys,
false,
- lowerBound,- upperBound);+ null,+ null);
Suggestion importance[1-10]: 4

__

Why: While the suggestion correctly identifies that RANK and DENSE_RANK disallow framing, the comment in the PR already acknowledges this ("Calcite rank operators disallow framing, so the ROWS/RANGE flag below is normalized away"). The current implementation passes lowerBound and upperBound which are likely normalized by Calcite. Explicitly passing null would be slightly clearer but is not critical since the normalization handles this.

Low
Suggestions up to commit 747d1b6
CategorySuggestion Impact
Possible issue
Validate ORDER BY clause presence

The RANK and DENSE_RANK window functions require an ORDER BY clause to function
correctly. Consider validating that orderKeys is not empty before creating the
aggregate call to prevent runtime errors or unexpected behavior when no ordering is
specified.

core/src/main/java/org/opensearch/sql/calcite/utils/PlanUtils.java [236-251]

 case RANK:
+ if (orderKeys.isEmpty()) {+ throw new IllegalArgumentException("RANK requires ORDER BY clause");+ }
return withOver(
context.relBuilder.aggregateCall(SqlStdOperatorTable.RANK),
partitions,
orderKeys,
false,
lowerBound,
upperBound);
case DENSE_RANK:
+ if (orderKeys.isEmpty()) {+ throw new IllegalArgumentException("DENSE_RANK requires ORDER BY clause");+ }
return withOver(
context.relBuilder.aggregateCall(SqlStdOperatorTable.DENSE_RANK),
partitions,
orderKeys,
false,
lowerBound,
upperBound);
Suggestion importance[1-10]: 5

__

Why: While RANK and DENSE_RANK do require an ORDER BY clause semantically, this validation may already be handled at the SQL parsing/validation layer. Adding redundant validation here could be useful for defensive programming, but without evidence of missing validation upstream, this is a moderate improvement for robustness rather than fixing a critical bug.

Low
Suggestions up to commit 352ffbd
CategorySuggestion Impact
General
Use Set for multiple condition checks

Consider using a Set or EnumSet for checking multiple function names instead of
chained OR conditions. This improves readability and makes it easier to add more
functions in the future.

core/src/main/java/org/opensearch/sql/calcite/CalciteRexNodeVisitor.java [775-777]

-if (functionName == BuiltinFunctionName.ROW_NUMBER- || functionName == BuiltinFunctionName.RANK- || functionName == BuiltinFunctionName.DENSE_RANK) {+private static final Set<BuiltinFunctionName> NO_FIELD_WINDOW_FUNCTIONS = + EnumSet.of(BuiltinFunctionName.ROW_NUMBER, BuiltinFunctionName.RANK, BuiltinFunctionName.DENSE_RANK);+if (NO_FIELD_WINDOW_FUNCTIONS.contains(functionName)) {+
Suggestion importance[1-10]: 4

__

Why: While using an EnumSet would improve maintainability and readability, the current chained OR condition with three items is still acceptable and clear. The suggestion is valid but offers only a moderate improvement in code style rather than fixing a critical issue.

Low

@dai-chen
dai-chenforce-pushed the support-rank-dense-rank-calcite branch from 352ffbd to 747d1b6CompareAugust 25, 2026 00:21
@github-actions

Copy link
Copy Markdown
Contributor

Persistent review updated to latest commit 747d1b6

@dai-chen
dai-chenforce-pushed the support-rank-dense-rank-calcite branch from 747d1b6 to 77ae6e0CompareAugust 25, 2026 18:17
@github-actions

Copy link
Copy Markdown
Contributor

Persistent review updated to latest commit 77ae6e0

Both were accepted by the SQL grammar and declared as
BuiltinFunctionName constants, but missing from WINDOW_FUNC_MAPPING, so
visitWindowFunction rejected them with "Unexpected window function".
Register both, extend the ROW_NUMBER bypass of aggregate signature
validation since they likewise take no arguments, and lower them to
SqlStdOperatorTable.RANK and DENSE_RANK.
WINDOW_FUNC_MAPPING is shared, so PPL eventstats/streamstats now resolve
these functions as well.
Related to opensearch-project#5168
Signed-off-by: Chen Dai <daichen@amazon.com>
@dai-chen
dai-chenforce-pushed the support-rank-dense-rank-calcite branch from 77ae6e0 to 544e467CompareAugust 26, 2026 22:13
@github-actions

Copy link
Copy Markdown
Contributor

Persistent review updated to latest commit 544e467

@dai-chendai-chen changed the title feat(sql): Support RANK/DENSE_RANK in unified SQLfeat(sql): Support RANK/DENSE_RANK for unified SQLAug 26, 2026
@dai-chen
dai-chen marked this pull request as ready for review August 26, 2026 23:10
gingeekrishna added a commit to gingeekrishna/sql that referenced this pull request Aug 28, 2026
WINDOW_FUNC_MAPPING (used by eventstats/streamstats) never supported
rank/dense_rank, but PPL's scalarWindowFunctionName grammar rule still
accepted the tokens, so `eventstats rank()` reached
CalciteRexNodeVisitor#visitWindowFunction and failed there with the
"not supported" message this PR improves. Meanwhile SQL's grammar
already accepts RANK()/DENSE_RANK() OVER (...), and opensearch-project#5720 is adding
real support for them on the SQL side via the same shared visitor.
Remove RANK/DENSE_RANK from scalarWindowFunctionName so PPL rejects
them at parse time instead of falling through to the shared
SQL/PPL visitor - this keeps the language separation at the parser
rather than relying on a WINDOW_FUNC_MAPPING check in shared planner
code, and avoids PPL silently gaining rank/dense_rank as a side effect
of opensearch-project#5720 registering them for SQL.
Updates the eventstats/streamstats tests added earlier in this PR to
expect a SyntaxCheckException (parse-time) instead of the semantic
"not supported" error, and switches the unrelated
visitWindowFunction-rejection unit test from rank() to percent_rank(),
which remains unsupported and still exercises that code path.
Signed-off-by: Radhakrishnan P <gingeekrishna@gmail.com>

@RyanL1997RyanL1997 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM. Just would like to double check do we need to add doc test/doc for this?

@RyanL1997

Copy link
Copy Markdown
Collaborator

In addtion, I found out that

The following isn't something you introduced in this PR. The existing ROW_NUMBER case has the same shape, and ROW_NUMBER() OVER (ORDER BY age) already returns 1 for every row. The new RANK/DENSE_RANK cases inherit it.

Behaviour:

SELECT age, RANK() OVER (ORDER BY age) FROM employees
actual → 4, 4, 4, 4 (the partition size)
expected → 1, 2, 3, 4

Why: the new cases forward the caller's lowerBound/upperBound unchanged, and for SQL those default to WindowFrame.rowsUnbounded() = UNBOUNDED PRECEDING … UNBOUNDED FOLLOWING, so every row is ranked over the whole partition. Calcite's own SqlToRelConverter.convertOver forces UNBOUNDED PRECEDING … CURRENT ROW for operators with allowsFraming() == false; this branch needs to do the same.

One knock-on: the plan assertions can't catch it. allowsFraming() only suppresses the frame in the printed digest, not in the executed Window.Group — so RANK() OVER (ORDER BY $2 NULLS FIRST) prints identically whichever frame was built. Confirming this needs an execution-level assertion.

@dai-chen
dai-chen merged commit c4f27b2 into opensearch-project:mainAug 28, 2026
42 checks passed
gingeekrishna added a commit to gingeekrishna/sql that referenced this pull request Aug 30, 2026
WINDOW_FUNC_MAPPING (used by eventstats/streamstats) never supported
rank/dense_rank, but PPL's scalarWindowFunctionName grammar rule still
accepted the tokens, so `eventstats rank()` reached
CalciteRexNodeVisitor#visitWindowFunction and failed there with the
"not supported" message this PR improves. Meanwhile SQL's grammar
already accepts RANK()/DENSE_RANK() OVER (...), and opensearch-project#5720 is adding
real support for them on the SQL side via the same shared visitor.
Remove RANK/DENSE_RANK from scalarWindowFunctionName so PPL rejects
them at parse time instead of falling through to the shared
SQL/PPL visitor - this keeps the language separation at the parser
rather than relying on a WINDOW_FUNC_MAPPING check in shared planner
code, and avoids PPL silently gaining rank/dense_rank as a side effect
of opensearch-project#5720 registering them for SQL.
Updates the eventstats/streamstats tests added earlier in this PR to
expect a SyntaxCheckException (parse-time) instead of the semantic
"not supported" error, and switches the unrelated
visitWindowFunction-rejection unit test from rank() to percent_rank(),
which remains unsupported and still exercises that code path.
Signed-off-by: Radhakrishnan P <gingeekrishna@gmail.com>
gingeekrishna added a commit to gingeekrishna/sql that referenced this pull request Sep 2, 2026
…layer
opensearch-project#5720 added rank/dense_rank to the WINDOW_FUNC_MAPPING shared by both
SQL's RANK()/DENSE_RANK() OVER (...) and PPL's eventstats/streamstats,
which silently enabled them for eventstats/streamstats too:
testRankingWindowFunctionsUnsupportedInEventstats/InStreamstats
(added earlier in this PR) started failing with "expected
ResponseException to be thrown, but nothing was thrown" once opensearch-project#5720
merged, since eventstats rank() now builds a real window call instead
of hitting the "not supported" check.
That's a real gap, not just a test artifact: PPL's eventstats/
streamstats grammar has no ORDER BY syntax at all, so ranking has no
defined ordering to rank by there - unlike SQL's OVER(), which at
least has (optional) ORDER BY in its own clause.
A prior commit on this branch tried fixing this by removing RANK/
DENSE_RANK from the PPL grammar entirely, rejecting them at parse
time. That broke a different, cross-repo contract: OpenSearch-
Dashboards' PPL linter (validated by the "PPL grammar compatibility"
CI check) expects `eventstats rank()` to parse successfully and be
flagged by a semantic-layer diagnostic instead, so it could no longer
produce that diagnostic once the query stopped parsing. That commit
was reverted.
Fix this at the semantic layer instead, where it belongs: PPL only
ever reaches CalciteRexNodeVisitor#visitWindowFunction through
eventstats/streamstats (no other PPL syntax builds a WindowFunction
node), so context.queryType == PPL is an exact, unambiguous signal for
"this is an eventstats/streamstats call". Filter rank/dense_rank out
of the WINDOW_FUNC_MAPPING lookup specifically when queryType is PPL,
so they fall through to the existing "not supported in eventstats/
streamstats" error - exactly the pre-opensearch-project#5720 behavior - while leaving
SQL's RANK()/DENSE_RANK() OVER (...) handling (and everything else)
completely untouched.
Signed-off-by: Radhakrishnan P <gingeekrishna@gmail.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

analytic-engineenhancementNew feature or requestSQL

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@dai-chen@RyanL1997
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(sql): Support RANK/DENSE_RANK for unified SQL - #5720

Merged
dai-chen merged 1 commit into
opensearch-project:mainfrom
dai-chen:support-rank-dense-rank-calcite
Aug 28, 2026
Merged

feat(sql): Support RANK/DENSE_RANK for unified SQL#5720
dai-chen merged 1 commit into
opensearch-project:mainfrom
dai-chen:support-rank-dense-rank-calcite

Conversation

@dai-chen

@dai-chendai-chen commented Aug 24, 2026

Copy link
Copy Markdown
Collaborator

Description

RANK() and DENSE_RANK() were already accepted by the SQL grammar and declared as built-in function names, but were never registered as window functions, so planning failed with Unexpected window function: RANK. This PR registers both and maps them to Calcite's standard RANK and DENSE_RANK operators. No grammar, lexer or AST changes were needed.

Related Issues

Part of #5248

Check List

  • New functionality includes testing.
  • New functionality has been documented.
  • New functionality has javadoc added.
  • New functionality has a user manual doc added.
  • New PPL command checklist all confirmed.
  • API changes companion pull request created.
  • Commits are signed per the DCO using --signoff or -s.
  • Public documentation issue/PR created.

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.
For more information on following Developer Certificate of Origin and signing off your commits, please check here.

@dai-chendai-chen self-assigned this Aug 24, 2026
@github-actions

github-actionsBot commented Aug 24, 2026

Copy link
Copy Markdown
Contributor

PR Reviewer Guide 🔍

(Review updated until commit 544e467)

Here are some key observations to aid the review process:

🧪 PR contains tests
🔒 No security concerns identified
✅ No TODO sections
🔀 No multiple PR themes
⚡ No major issues detected

@github-actions

github-actionsBot commented Aug 24, 2026

Copy link
Copy Markdown
Contributor

PR Code Suggestions ✨

Latest suggestions up to 544e467

Explore these optional code suggestions:

CategorySuggestion Impact
General
Explicitly pass null for frame bounds

Since RANK and DENSE_RANK disallow framing, passing lowerBound and upperBound
parameters may cause unexpected behavior or errors. Consider passing null for both
bounds to explicitly indicate no framing is used.

core/src/main/java/org/opensearch/sql/calcite/utils/PlanUtils.java [236-251]

 case RANK:
return withOver(
context.relBuilder.aggregateCall(SqlStdOperatorTable.RANK),
partitions,
orderKeys,
false,
- lowerBound,- upperBound);+ null,+ null);
case DENSE_RANK:
return withOver(
context.relBuilder.aggregateCall(SqlStdOperatorTable.DENSE_RANK),
partitions,
orderKeys,
false,
- lowerBound,- upperBound);+ null,+ null);
Suggestion importance[1-10]: 7

__

Why: The suggestion correctly identifies that RANK and DENSE_RANK disallow framing, and the comment in the code confirms this. Passing null instead of lowerBound and upperBound would make the intent more explicit and prevent potential issues, though the current implementation may already handle this correctly through normalization.

Medium

Previous suggestions

Suggestions up to commit 77ae6e0
CategorySuggestion Impact
Possible issue
Explicitly pass null for frame bounds

Since RANK and DENSE_RANK disallow framing, passing lowerBound and upperBound
parameters may cause unexpected behavior or errors. Consider explicitly passing null
for both bounds to ensure proper handling of these ranking functions.

core/src/main/java/org/opensearch/sql/calcite/utils/PlanUtils.java [236-251]

 case RANK:
return withOver(
context.relBuilder.aggregateCall(SqlStdOperatorTable.RANK),
partitions,
orderKeys,
false,
- lowerBound,- upperBound);+ null,+ null);
case DENSE_RANK:
return withOver(
context.relBuilder.aggregateCall(SqlStdOperatorTable.DENSE_RANK),
partitions,
orderKeys,
false,
- lowerBound,- upperBound);+ null,+ null);
Suggestion importance[1-10]: 4

__

Why: While the suggestion correctly identifies that RANK and DENSE_RANK disallow framing, the comment in the PR already acknowledges this ("Calcite rank operators disallow framing, so the ROWS/RANGE flag below is normalized away"). The current implementation passes lowerBound and upperBound which are likely normalized by Calcite. Explicitly passing null would be slightly clearer but is not critical since the normalization handles this.

Low
Suggestions up to commit 747d1b6
CategorySuggestion Impact
Possible issue
Validate ORDER BY clause presence

The RANK and DENSE_RANK window functions require an ORDER BY clause to function
correctly. Consider validating that orderKeys is not empty before creating the
aggregate call to prevent runtime errors or unexpected behavior when no ordering is
specified.

core/src/main/java/org/opensearch/sql/calcite/utils/PlanUtils.java [236-251]

 case RANK:
+ if (orderKeys.isEmpty()) {+ throw new IllegalArgumentException("RANK requires ORDER BY clause");+ }
return withOver(
context.relBuilder.aggregateCall(SqlStdOperatorTable.RANK),
partitions,
orderKeys,
false,
lowerBound,
upperBound);
case DENSE_RANK:
+ if (orderKeys.isEmpty()) {+ throw new IllegalArgumentException("DENSE_RANK requires ORDER BY clause");+ }
return withOver(
context.relBuilder.aggregateCall(SqlStdOperatorTable.DENSE_RANK),
partitions,
orderKeys,
false,
lowerBound,
upperBound);
Suggestion importance[1-10]: 5

__

Why: While RANK and DENSE_RANK do require an ORDER BY clause semantically, this validation may already be handled at the SQL parsing/validation layer. Adding redundant validation here could be useful for defensive programming, but without evidence of missing validation upstream, this is a moderate improvement for robustness rather than fixing a critical bug.

Low
Suggestions up to commit 352ffbd
CategorySuggestion Impact
General
Use Set for multiple condition checks

Consider using a Set or EnumSet for checking multiple function names instead of
chained OR conditions. This improves readability and makes it easier to add more
functions in the future.

core/src/main/java/org/opensearch/sql/calcite/CalciteRexNodeVisitor.java [775-777]

-if (functionName == BuiltinFunctionName.ROW_NUMBER- || functionName == BuiltinFunctionName.RANK- || functionName == BuiltinFunctionName.DENSE_RANK) {+private static final Set<BuiltinFunctionName> NO_FIELD_WINDOW_FUNCTIONS = + EnumSet.of(BuiltinFunctionName.ROW_NUMBER, BuiltinFunctionName.RANK, BuiltinFunctionName.DENSE_RANK);+if (NO_FIELD_WINDOW_FUNCTIONS.contains(functionName)) {+
Suggestion importance[1-10]: 4

__

Why: While using an EnumSet would improve maintainability and readability, the current chained OR condition with three items is still acceptable and clear. The suggestion is valid but offers only a moderate improvement in code style rather than fixing a critical issue.

Low

@dai-chen
dai-chenforce-pushed the support-rank-dense-rank-calcite branch from 352ffbd to 747d1b6CompareAugust 25, 2026 00:21
@github-actions

Copy link
Copy Markdown
Contributor

Persistent review updated to latest commit 747d1b6

@dai-chen
dai-chenforce-pushed the support-rank-dense-rank-calcite branch from 747d1b6 to 77ae6e0CompareAugust 25, 2026 18:17
@github-actions

Copy link
Copy Markdown
Contributor

Persistent review updated to latest commit 77ae6e0

Both were accepted by the SQL grammar and declared as
BuiltinFunctionName constants, but missing from WINDOW_FUNC_MAPPING, so
visitWindowFunction rejected them with "Unexpected window function".
Register both, extend the ROW_NUMBER bypass of aggregate signature
validation since they likewise take no arguments, and lower them to
SqlStdOperatorTable.RANK and DENSE_RANK.
WINDOW_FUNC_MAPPING is shared, so PPL eventstats/streamstats now resolve
these functions as well.
Related to opensearch-project#5168
Signed-off-by: Chen Dai <daichen@amazon.com>
@dai-chen
dai-chenforce-pushed the support-rank-dense-rank-calcite branch from 77ae6e0 to 544e467CompareAugust 26, 2026 22:13
@github-actions

Copy link
Copy Markdown
Contributor

Persistent review updated to latest commit 544e467

@dai-chendai-chen changed the title feat(sql): Support RANK/DENSE_RANK in unified SQLfeat(sql): Support RANK/DENSE_RANK for unified SQLAug 26, 2026
@dai-chen
dai-chen marked this pull request as ready for review August 26, 2026 23:10
gingeekrishna added a commit to gingeekrishna/sql that referenced this pull request Aug 28, 2026
WINDOW_FUNC_MAPPING (used by eventstats/streamstats) never supported
rank/dense_rank, but PPL's scalarWindowFunctionName grammar rule still
accepted the tokens, so `eventstats rank()` reached
CalciteRexNodeVisitor#visitWindowFunction and failed there with the
"not supported" message this PR improves. Meanwhile SQL's grammar
already accepts RANK()/DENSE_RANK() OVER (...), and opensearch-project#5720 is adding
real support for them on the SQL side via the same shared visitor.
Remove RANK/DENSE_RANK from scalarWindowFunctionName so PPL rejects
them at parse time instead of falling through to the shared
SQL/PPL visitor - this keeps the language separation at the parser
rather than relying on a WINDOW_FUNC_MAPPING check in shared planner
code, and avoids PPL silently gaining rank/dense_rank as a side effect
of opensearch-project#5720 registering them for SQL.
Updates the eventstats/streamstats tests added earlier in this PR to
expect a SyntaxCheckException (parse-time) instead of the semantic
"not supported" error, and switches the unrelated
visitWindowFunction-rejection unit test from rank() to percent_rank(),
which remains unsupported and still exercises that code path.
Signed-off-by: Radhakrishnan P <gingeekrishna@gmail.com>

@RyanL1997RyanL1997 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM. Just would like to double check do we need to add doc test/doc for this?

@RyanL1997

Copy link
Copy Markdown
Collaborator

In addtion, I found out that

The following isn't something you introduced in this PR. The existing ROW_NUMBER case has the same shape, and ROW_NUMBER() OVER (ORDER BY age) already returns 1 for every row. The new RANK/DENSE_RANK cases inherit it.

Behaviour:

SELECT age, RANK() OVER (ORDER BY age) FROM employees
actual → 4, 4, 4, 4 (the partition size)
expected → 1, 2, 3, 4

Why: the new cases forward the caller's lowerBound/upperBound unchanged, and for SQL those default to WindowFrame.rowsUnbounded() = UNBOUNDED PRECEDING … UNBOUNDED FOLLOWING, so every row is ranked over the whole partition. Calcite's own SqlToRelConverter.convertOver forces UNBOUNDED PRECEDING … CURRENT ROW for operators with allowsFraming() == false; this branch needs to do the same.

One knock-on: the plan assertions can't catch it. allowsFraming() only suppresses the frame in the printed digest, not in the executed Window.Group — so RANK() OVER (ORDER BY $2 NULLS FIRST) prints identically whichever frame was built. Confirming this needs an execution-level assertion.

@dai-chen
dai-chen merged commit c4f27b2 into opensearch-project:mainAug 28, 2026
42 checks passed
gingeekrishna added a commit to gingeekrishna/sql that referenced this pull request Aug 30, 2026
WINDOW_FUNC_MAPPING (used by eventstats/streamstats) never supported
rank/dense_rank, but PPL's scalarWindowFunctionName grammar rule still
accepted the tokens, so `eventstats rank()` reached
CalciteRexNodeVisitor#visitWindowFunction and failed there with the
"not supported" message this PR improves. Meanwhile SQL's grammar
already accepts RANK()/DENSE_RANK() OVER (...), and opensearch-project#5720 is adding
real support for them on the SQL side via the same shared visitor.
Remove RANK/DENSE_RANK from scalarWindowFunctionName so PPL rejects
them at parse time instead of falling through to the shared
SQL/PPL visitor - this keeps the language separation at the parser
rather than relying on a WINDOW_FUNC_MAPPING check in shared planner
code, and avoids PPL silently gaining rank/dense_rank as a side effect
of opensearch-project#5720 registering them for SQL.
Updates the eventstats/streamstats tests added earlier in this PR to
expect a SyntaxCheckException (parse-time) instead of the semantic
"not supported" error, and switches the unrelated
visitWindowFunction-rejection unit test from rank() to percent_rank(),
which remains unsupported and still exercises that code path.
Signed-off-by: Radhakrishnan P <gingeekrishna@gmail.com>
gingeekrishna added a commit to gingeekrishna/sql that referenced this pull request Sep 2, 2026
…layer
opensearch-project#5720 added rank/dense_rank to the WINDOW_FUNC_MAPPING shared by both
SQL's RANK()/DENSE_RANK() OVER (...) and PPL's eventstats/streamstats,
which silently enabled them for eventstats/streamstats too:
testRankingWindowFunctionsUnsupportedInEventstats/InStreamstats
(added earlier in this PR) started failing with "expected
ResponseException to be thrown, but nothing was thrown" once opensearch-project#5720
merged, since eventstats rank() now builds a real window call instead
of hitting the "not supported" check.
That's a real gap, not just a test artifact: PPL's eventstats/
streamstats grammar has no ORDER BY syntax at all, so ranking has no
defined ordering to rank by there - unlike SQL's OVER(), which at
least has (optional) ORDER BY in its own clause.
A prior commit on this branch tried fixing this by removing RANK/
DENSE_RANK from the PPL grammar entirely, rejecting them at parse
time. That broke a different, cross-repo contract: OpenSearch-
Dashboards' PPL linter (validated by the "PPL grammar compatibility"
CI check) expects `eventstats rank()` to parse successfully and be
flagged by a semantic-layer diagnostic instead, so it could no longer
produce that diagnostic once the query stopped parsing. That commit
was reverted.
Fix this at the semantic layer instead, where it belongs: PPL only
ever reaches CalciteRexNodeVisitor#visitWindowFunction through
eventstats/streamstats (no other PPL syntax builds a WindowFunction
node), so context.queryType == PPL is an exact, unambiguous signal for
"this is an eventstats/streamstats call". Filter rank/dense_rank out
of the WINDOW_FUNC_MAPPING lookup specifically when queryType is PPL,
so they fall through to the existing "not supported in eventstats/
streamstats" error - exactly the pre-opensearch-project#5720 behavior - while leaving
SQL's RANK()/DENSE_RANK() OVER (...) handling (and everything else)
completely untouched.
Signed-off-by: Radhakrishnan P <gingeekrishna@gmail.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

analytic-engineenhancementNew feature or requestSQL

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@dai-chen@RyanL1997
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(sql): Support RANK/DENSE_RANK for unified SQL - #5720

Merged
dai-chen merged 1 commit into
opensearch-project:mainfrom
dai-chen:support-rank-dense-rank-calcite
Aug 28, 2026
Merged

feat(sql): Support RANK/DENSE_RANK for unified SQL#5720
dai-chen merged 1 commit into
opensearch-project:mainfrom
dai-chen:support-rank-dense-rank-calcite

Conversation

@dai-chen

@dai-chendai-chen commented Aug 24, 2026

Copy link
Copy Markdown
Collaborator

Description

RANK() and DENSE_RANK() were already accepted by the SQL grammar and declared as built-in function names, but were never registered as window functions, so planning failed with Unexpected window function: RANK. This PR registers both and maps them to Calcite's standard RANK and DENSE_RANK operators. No grammar, lexer or AST changes were needed.

Related Issues

Part of #5248

Check List

  • New functionality includes testing.
  • New functionality has been documented.
  • New functionality has javadoc added.
  • New functionality has a user manual doc added.
  • New PPL command checklist all confirmed.
  • API changes companion pull request created.
  • Commits are signed per the DCO using --signoff or -s.
  • Public documentation issue/PR created.

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.
For more information on following Developer Certificate of Origin and signing off your commits, please check here.

@dai-chendai-chen self-assigned this Aug 24, 2026
@github-actions

github-actionsBot commented Aug 24, 2026

Copy link
Copy Markdown
Contributor

PR Reviewer Guide 🔍

(Review updated until commit 544e467)

Here are some key observations to aid the review process:

🧪 PR contains tests
🔒 No security concerns identified
✅ No TODO sections
🔀 No multiple PR themes
⚡ No major issues detected

@github-actions

github-actionsBot commented Aug 24, 2026

Copy link
Copy Markdown
Contributor

PR Code Suggestions ✨

Latest suggestions up to 544e467

Explore these optional code suggestions:

CategorySuggestion Impact
General
Explicitly pass null for frame bounds

Since RANK and DENSE_RANK disallow framing, passing lowerBound and upperBound
parameters may cause unexpected behavior or errors. Consider passing null for both
bounds to explicitly indicate no framing is used.

core/src/main/java/org/opensearch/sql/calcite/utils/PlanUtils.java [236-251]

 case RANK:
return withOver(
context.relBuilder.aggregateCall(SqlStdOperatorTable.RANK),
partitions,
orderKeys,
false,
- lowerBound,- upperBound);+ null,+ null);
case DENSE_RANK:
return withOver(
context.relBuilder.aggregateCall(SqlStdOperatorTable.DENSE_RANK),
partitions,
orderKeys,
false,
- lowerBound,- upperBound);+ null,+ null);
Suggestion importance[1-10]: 7

__

Why: The suggestion correctly identifies that RANK and DENSE_RANK disallow framing, and the comment in the code confirms this. Passing null instead of lowerBound and upperBound would make the intent more explicit and prevent potential issues, though the current implementation may already handle this correctly through normalization.

Medium

Previous suggestions

Suggestions up to commit 77ae6e0
CategorySuggestion Impact
Possible issue
Explicitly pass null for frame bounds

Since RANK and DENSE_RANK disallow framing, passing lowerBound and upperBound
parameters may cause unexpected behavior or errors. Consider explicitly passing null
for both bounds to ensure proper handling of these ranking functions.

core/src/main/java/org/opensearch/sql/calcite/utils/PlanUtils.java [236-251]

 case RANK:
return withOver(
context.relBuilder.aggregateCall(SqlStdOperatorTable.RANK),
partitions,
orderKeys,
false,
- lowerBound,- upperBound);+ null,+ null);
case DENSE_RANK:
return withOver(
context.relBuilder.aggregateCall(SqlStdOperatorTable.DENSE_RANK),
partitions,
orderKeys,
false,
- lowerBound,- upperBound);+ null,+ null);
Suggestion importance[1-10]: 4

__

Why: While the suggestion correctly identifies that RANK and DENSE_RANK disallow framing, the comment in the PR already acknowledges this ("Calcite rank operators disallow framing, so the ROWS/RANGE flag below is normalized away"). The current implementation passes lowerBound and upperBound which are likely normalized by Calcite. Explicitly passing null would be slightly clearer but is not critical since the normalization handles this.

Low
Suggestions up to commit 747d1b6
CategorySuggestion Impact
Possible issue
Validate ORDER BY clause presence

The RANK and DENSE_RANK window functions require an ORDER BY clause to function
correctly. Consider validating that orderKeys is not empty before creating the
aggregate call to prevent runtime errors or unexpected behavior when no ordering is
specified.

core/src/main/java/org/opensearch/sql/calcite/utils/PlanUtils.java [236-251]

 case RANK:
+ if (orderKeys.isEmpty()) {+ throw new IllegalArgumentException("RANK requires ORDER BY clause");+ }
return withOver(
context.relBuilder.aggregateCall(SqlStdOperatorTable.RANK),
partitions,
orderKeys,
false,
lowerBound,
upperBound);
case DENSE_RANK:
+ if (orderKeys.isEmpty()) {+ throw new IllegalArgumentException("DENSE_RANK requires ORDER BY clause");+ }
return withOver(
context.relBuilder.aggregateCall(SqlStdOperatorTable.DENSE_RANK),
partitions,
orderKeys,
false,
lowerBound,
upperBound);
Suggestion importance[1-10]: 5

__

Why: While RANK and DENSE_RANK do require an ORDER BY clause semantically, this validation may already be handled at the SQL parsing/validation layer. Adding redundant validation here could be useful for defensive programming, but without evidence of missing validation upstream, this is a moderate improvement for robustness rather than fixing a critical bug.

Low
Suggestions up to commit 352ffbd
CategorySuggestion Impact
General
Use Set for multiple condition checks

Consider using a Set or EnumSet for checking multiple function names instead of
chained OR conditions. This improves readability and makes it easier to add more
functions in the future.

core/src/main/java/org/opensearch/sql/calcite/CalciteRexNodeVisitor.java [775-777]

-if (functionName == BuiltinFunctionName.ROW_NUMBER- || functionName == BuiltinFunctionName.RANK- || functionName == BuiltinFunctionName.DENSE_RANK) {+private static final Set<BuiltinFunctionName> NO_FIELD_WINDOW_FUNCTIONS = + EnumSet.of(BuiltinFunctionName.ROW_NUMBER, BuiltinFunctionName.RANK, BuiltinFunctionName.DENSE_RANK);+if (NO_FIELD_WINDOW_FUNCTIONS.contains(functionName)) {+
Suggestion importance[1-10]: 4

__

Why: While using an EnumSet would improve maintainability and readability, the current chained OR condition with three items is still acceptable and clear. The suggestion is valid but offers only a moderate improvement in code style rather than fixing a critical issue.

Low

@dai-chen
dai-chenforce-pushed the support-rank-dense-rank-calcite branch from 352ffbd to 747d1b6CompareAugust 25, 2026 00:21
@github-actions

Copy link
Copy Markdown
Contributor

Persistent review updated to latest commit 747d1b6

@dai-chen
dai-chenforce-pushed the support-rank-dense-rank-calcite branch from 747d1b6 to 77ae6e0CompareAugust 25, 2026 18:17
@github-actions

Copy link
Copy Markdown
Contributor

Persistent review updated to latest commit 77ae6e0

Both were accepted by the SQL grammar and declared as
BuiltinFunctionName constants, but missing from WINDOW_FUNC_MAPPING, so
visitWindowFunction rejected them with "Unexpected window function".
Register both, extend the ROW_NUMBER bypass of aggregate signature
validation since they likewise take no arguments, and lower them to
SqlStdOperatorTable.RANK and DENSE_RANK.
WINDOW_FUNC_MAPPING is shared, so PPL eventstats/streamstats now resolve
these functions as well.
Related to opensearch-project#5168
Signed-off-by: Chen Dai <daichen@amazon.com>
@dai-chen
dai-chenforce-pushed the support-rank-dense-rank-calcite branch from 77ae6e0 to 544e467CompareAugust 26, 2026 22:13
@github-actions

Copy link
Copy Markdown
Contributor

Persistent review updated to latest commit 544e467

@dai-chendai-chen changed the title feat(sql): Support RANK/DENSE_RANK in unified SQLfeat(sql): Support RANK/DENSE_RANK for unified SQLAug 26, 2026
@dai-chen
dai-chen marked this pull request as ready for review August 26, 2026 23:10
gingeekrishna added a commit to gingeekrishna/sql that referenced this pull request Aug 28, 2026
WINDOW_FUNC_MAPPING (used by eventstats/streamstats) never supported
rank/dense_rank, but PPL's scalarWindowFunctionName grammar rule still
accepted the tokens, so `eventstats rank()` reached
CalciteRexNodeVisitor#visitWindowFunction and failed there with the
"not supported" message this PR improves. Meanwhile SQL's grammar
already accepts RANK()/DENSE_RANK() OVER (...), and opensearch-project#5720 is adding
real support for them on the SQL side via the same shared visitor.
Remove RANK/DENSE_RANK from scalarWindowFunctionName so PPL rejects
them at parse time instead of falling through to the shared
SQL/PPL visitor - this keeps the language separation at the parser
rather than relying on a WINDOW_FUNC_MAPPING check in shared planner
code, and avoids PPL silently gaining rank/dense_rank as a side effect
of opensearch-project#5720 registering them for SQL.
Updates the eventstats/streamstats tests added earlier in this PR to
expect a SyntaxCheckException (parse-time) instead of the semantic
"not supported" error, and switches the unrelated
visitWindowFunction-rejection unit test from rank() to percent_rank(),
which remains unsupported and still exercises that code path.
Signed-off-by: Radhakrishnan P <gingeekrishna@gmail.com>

@RyanL1997RyanL1997 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM. Just would like to double check do we need to add doc test/doc for this?

@RyanL1997

Copy link
Copy Markdown
Collaborator

In addtion, I found out that

The following isn't something you introduced in this PR. The existing ROW_NUMBER case has the same shape, and ROW_NUMBER() OVER (ORDER BY age) already returns 1 for every row. The new RANK/DENSE_RANK cases inherit it.

Behaviour:

SELECT age, RANK() OVER (ORDER BY age) FROM employees
actual → 4, 4, 4, 4 (the partition size)
expected → 1, 2, 3, 4

Why: the new cases forward the caller's lowerBound/upperBound unchanged, and for SQL those default to WindowFrame.rowsUnbounded() = UNBOUNDED PRECEDING … UNBOUNDED FOLLOWING, so every row is ranked over the whole partition. Calcite's own SqlToRelConverter.convertOver forces UNBOUNDED PRECEDING … CURRENT ROW for operators with allowsFraming() == false; this branch needs to do the same.

One knock-on: the plan assertions can't catch it. allowsFraming() only suppresses the frame in the printed digest, not in the executed Window.Group — so RANK() OVER (ORDER BY $2 NULLS FIRST) prints identically whichever frame was built. Confirming this needs an execution-level assertion.

@dai-chen
dai-chen merged commit c4f27b2 into opensearch-project:mainAug 28, 2026
42 checks passed
gingeekrishna added a commit to gingeekrishna/sql that referenced this pull request Aug 30, 2026
WINDOW_FUNC_MAPPING (used by eventstats/streamstats) never supported
rank/dense_rank, but PPL's scalarWindowFunctionName grammar rule still
accepted the tokens, so `eventstats rank()` reached
CalciteRexNodeVisitor#visitWindowFunction and failed there with the
"not supported" message this PR improves. Meanwhile SQL's grammar
already accepts RANK()/DENSE_RANK() OVER (...), and opensearch-project#5720 is adding
real support for them on the SQL side via the same shared visitor.
Remove RANK/DENSE_RANK from scalarWindowFunctionName so PPL rejects
them at parse time instead of falling through to the shared
SQL/PPL visitor - this keeps the language separation at the parser
rather than relying on a WINDOW_FUNC_MAPPING check in shared planner
code, and avoids PPL silently gaining rank/dense_rank as a side effect
of opensearch-project#5720 registering them for SQL.
Updates the eventstats/streamstats tests added earlier in this PR to
expect a SyntaxCheckException (parse-time) instead of the semantic
"not supported" error, and switches the unrelated
visitWindowFunction-rejection unit test from rank() to percent_rank(),
which remains unsupported and still exercises that code path.
Signed-off-by: Radhakrishnan P <gingeekrishna@gmail.com>
gingeekrishna added a commit to gingeekrishna/sql that referenced this pull request Sep 2, 2026
…layer
opensearch-project#5720 added rank/dense_rank to the WINDOW_FUNC_MAPPING shared by both
SQL's RANK()/DENSE_RANK() OVER (...) and PPL's eventstats/streamstats,
which silently enabled them for eventstats/streamstats too:
testRankingWindowFunctionsUnsupportedInEventstats/InStreamstats
(added earlier in this PR) started failing with "expected
ResponseException to be thrown, but nothing was thrown" once opensearch-project#5720
merged, since eventstats rank() now builds a real window call instead
of hitting the "not supported" check.
That's a real gap, not just a test artifact: PPL's eventstats/
streamstats grammar has no ORDER BY syntax at all, so ranking has no
defined ordering to rank by there - unlike SQL's OVER(), which at
least has (optional) ORDER BY in its own clause.
A prior commit on this branch tried fixing this by removing RANK/
DENSE_RANK from the PPL grammar entirely, rejecting them at parse
time. That broke a different, cross-repo contract: OpenSearch-
Dashboards' PPL linter (validated by the "PPL grammar compatibility"
CI check) expects `eventstats rank()` to parse successfully and be
flagged by a semantic-layer diagnostic instead, so it could no longer
produce that diagnostic once the query stopped parsing. That commit
was reverted.
Fix this at the semantic layer instead, where it belongs: PPL only
ever reaches CalciteRexNodeVisitor#visitWindowFunction through
eventstats/streamstats (no other PPL syntax builds a WindowFunction
node), so context.queryType == PPL is an exact, unambiguous signal for
"this is an eventstats/streamstats call". Filter rank/dense_rank out
of the WINDOW_FUNC_MAPPING lookup specifically when queryType is PPL,
so they fall through to the existing "not supported in eventstats/
streamstats" error - exactly the pre-opensearch-project#5720 behavior - while leaving
SQL's RANK()/DENSE_RANK() OVER (...) handling (and everything else)
completely untouched.
Signed-off-by: Radhakrishnan P <gingeekrishna@gmail.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

analytic-engineenhancementNew feature or requestSQL

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@dai-chen@RyanL1997
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

feat(sql): Support RANK/DENSE_RANK for unified SQL - #5720

Merged
dai-chen merged 1 commit into
opensearch-project:mainfrom
dai-chen:support-rank-dense-rank-calcite
Aug 28, 2026
Merged

feat(sql): Support RANK/DENSE_RANK for unified SQL#5720
dai-chen merged 1 commit into
opensearch-project:mainfrom
dai-chen:support-rank-dense-rank-calcite

Conversation

@dai-chen

@dai-chendai-chen commented Aug 24, 2026

Copy link
Copy Markdown
Collaborator

Description

RANK() and DENSE_RANK() were already accepted by the SQL grammar and declared as built-in function names, but were never registered as window functions, so planning failed with Unexpected window function: RANK. This PR registers both and maps them to Calcite's standard RANK and DENSE_RANK operators. No grammar, lexer or AST changes were needed.

Related Issues

Part of #5248

Check List

  • New functionality includes testing.
  • New functionality has been documented.
  • New functionality has javadoc added.
  • New functionality has a user manual doc added.
  • New PPL command checklist all confirmed.
  • API changes companion pull request created.
  • Commits are signed per the DCO using --signoff or -s.
  • Public documentation issue/PR created.

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.
For more information on following Developer Certificate of Origin and signing off your commits, please check here.

@dai-chendai-chen self-assigned this Aug 24, 2026
@github-actions

github-actionsBot commented Aug 24, 2026

Copy link
Copy Markdown
Contributor

PR Reviewer Guide 🔍

(Review updated until commit 544e467)

Here are some key observations to aid the review process:

🧪 PR contains tests
🔒 No security concerns identified
✅ No TODO sections
🔀 No multiple PR themes
⚡ No major issues detected

@github-actions

github-actionsBot commented Aug 24, 2026

Copy link
Copy Markdown
Contributor

PR Code Suggestions ✨

Latest suggestions up to 544e467

Explore these optional code suggestions:

CategorySuggestion Impact
General
Explicitly pass null for frame bounds

Since RANK and DENSE_RANK disallow framing, passing lowerBound and upperBound
parameters may cause unexpected behavior or errors. Consider passing null for both
bounds to explicitly indicate no framing is used.

core/src/main/java/org/opensearch/sql/calcite/utils/PlanUtils.java [236-251]

 case RANK:
return withOver(
context.relBuilder.aggregateCall(SqlStdOperatorTable.RANK),
partitions,
orderKeys,
false,
- lowerBound,- upperBound);+ null,+ null);
case DENSE_RANK:
return withOver(
context.relBuilder.aggregateCall(SqlStdOperatorTable.DENSE_RANK),
partitions,
orderKeys,
false,
- lowerBound,- upperBound);+ null,+ null);
Suggestion importance[1-10]: 7

__

Why: The suggestion correctly identifies that RANK and DENSE_RANK disallow framing, and the comment in the code confirms this. Passing null instead of lowerBound and upperBound would make the intent more explicit and prevent potential issues, though the current implementation may already handle this correctly through normalization.

Medium

Previous suggestions

Suggestions up to commit 77ae6e0
CategorySuggestion Impact
Possible issue
Explicitly pass null for frame bounds

Since RANK and DENSE_RANK disallow framing, passing lowerBound and upperBound
parameters may cause unexpected behavior or errors. Consider explicitly passing null
for both bounds to ensure proper handling of these ranking functions.

core/src/main/java/org/opensearch/sql/calcite/utils/PlanUtils.java [236-251]

 case RANK:
return withOver(
context.relBuilder.aggregateCall(SqlStdOperatorTable.RANK),
partitions,
orderKeys,
false,
- lowerBound,- upperBound);+ null,+ null);
case DENSE_RANK:
return withOver(
context.relBuilder.aggregateCall(SqlStdOperatorTable.DENSE_RANK),
partitions,
orderKeys,
false,
- lowerBound,- upperBound);+ null,+ null);
Suggestion importance[1-10]: 4

__

Why: While the suggestion correctly identifies that RANK and DENSE_RANK disallow framing, the comment in the PR already acknowledges this ("Calcite rank operators disallow framing, so the ROWS/RANGE flag below is normalized away"). The current implementation passes lowerBound and upperBound which are likely normalized by Calcite. Explicitly passing null would be slightly clearer but is not critical since the normalization handles this.

Low
Suggestions up to commit 747d1b6
CategorySuggestion Impact
Possible issue
Validate ORDER BY clause presence

The RANK and DENSE_RANK window functions require an ORDER BY clause to function
correctly. Consider validating that orderKeys is not empty before creating the
aggregate call to prevent runtime errors or unexpected behavior when no ordering is
specified.

core/src/main/java/org/opensearch/sql/calcite/utils/PlanUtils.java [236-251]

 case RANK:
+ if (orderKeys.isEmpty()) {+ throw new IllegalArgumentException("RANK requires ORDER BY clause");+ }
return withOver(
context.relBuilder.aggregateCall(SqlStdOperatorTable.RANK),
partitions,
orderKeys,
false,
lowerBound,
upperBound);
case DENSE_RANK:
+ if (orderKeys.isEmpty()) {+ throw new IllegalArgumentException("DENSE_RANK requires ORDER BY clause");+ }
return withOver(
context.relBuilder.aggregateCall(SqlStdOperatorTable.DENSE_RANK),
partitions,
orderKeys,
false,
lowerBound,
upperBound);
Suggestion importance[1-10]: 5

__

Why: While RANK and DENSE_RANK do require an ORDER BY clause semantically, this validation may already be handled at the SQL parsing/validation layer. Adding redundant validation here could be useful for defensive programming, but without evidence of missing validation upstream, this is a moderate improvement for robustness rather than fixing a critical bug.

Low
Suggestions up to commit 352ffbd
CategorySuggestion Impact
General
Use Set for multiple condition checks

Consider using a Set or EnumSet for checking multiple function names instead of
chained OR conditions. This improves readability and makes it easier to add more
functions in the future.

core/src/main/java/org/opensearch/sql/calcite/CalciteRexNodeVisitor.java [775-777]

-if (functionName == BuiltinFunctionName.ROW_NUMBER- || functionName == BuiltinFunctionName.RANK- || functionName == BuiltinFunctionName.DENSE_RANK) {+private static final Set<BuiltinFunctionName> NO_FIELD_WINDOW_FUNCTIONS = + EnumSet.of(BuiltinFunctionName.ROW_NUMBER, BuiltinFunctionName.RANK, BuiltinFunctionName.DENSE_RANK);+if (NO_FIELD_WINDOW_FUNCTIONS.contains(functionName)) {+
Suggestion importance[1-10]: 4

__

Why: While using an EnumSet would improve maintainability and readability, the current chained OR condition with three items is still acceptable and clear. The suggestion is valid but offers only a moderate improvement in code style rather than fixing a critical issue.

Low

@dai-chen
dai-chenforce-pushed the support-rank-dense-rank-calcite branch from 352ffbd to 747d1b6CompareAugust 25, 2026 00:21
@github-actions

Copy link
Copy Markdown
Contributor

Persistent review updated to latest commit 747d1b6

@dai-chen
dai-chenforce-pushed the support-rank-dense-rank-calcite branch from 747d1b6 to 77ae6e0CompareAugust 25, 2026 18:17
@github-actions

Copy link
Copy Markdown
Contributor

Persistent review updated to latest commit 77ae6e0

Both were accepted by the SQL grammar and declared as
BuiltinFunctionName constants, but missing from WINDOW_FUNC_MAPPING, so
visitWindowFunction rejected them with "Unexpected window function".
Register both, extend the ROW_NUMBER bypass of aggregate signature
validation since they likewise take no arguments, and lower them to
SqlStdOperatorTable.RANK and DENSE_RANK.
WINDOW_FUNC_MAPPING is shared, so PPL eventstats/streamstats now resolve
these functions as well.
Related to opensearch-project#5168
Signed-off-by: Chen Dai <daichen@amazon.com>
@dai-chen
dai-chenforce-pushed the support-rank-dense-rank-calcite branch from 77ae6e0 to 544e467CompareAugust 26, 2026 22:13
@github-actions

Copy link
Copy Markdown
Contributor

Persistent review updated to latest commit 544e467

@dai-chendai-chen changed the title feat(sql): Support RANK/DENSE_RANK in unified SQLfeat(sql): Support RANK/DENSE_RANK for unified SQLAug 26, 2026
@dai-chen
dai-chen marked this pull request as ready for review August 26, 2026 23:10
gingeekrishna added a commit to gingeekrishna/sql that referenced this pull request Aug 28, 2026
WINDOW_FUNC_MAPPING (used by eventstats/streamstats) never supported
rank/dense_rank, but PPL's scalarWindowFunctionName grammar rule still
accepted the tokens, so `eventstats rank()` reached
CalciteRexNodeVisitor#visitWindowFunction and failed there with the
"not supported" message this PR improves. Meanwhile SQL's grammar
already accepts RANK()/DENSE_RANK() OVER (...), and opensearch-project#5720 is adding
real support for them on the SQL side via the same shared visitor.
Remove RANK/DENSE_RANK from scalarWindowFunctionName so PPL rejects
them at parse time instead of falling through to the shared
SQL/PPL visitor - this keeps the language separation at the parser
rather than relying on a WINDOW_FUNC_MAPPING check in shared planner
code, and avoids PPL silently gaining rank/dense_rank as a side effect
of opensearch-project#5720 registering them for SQL.
Updates the eventstats/streamstats tests added earlier in this PR to
expect a SyntaxCheckException (parse-time) instead of the semantic
"not supported" error, and switches the unrelated
visitWindowFunction-rejection unit test from rank() to percent_rank(),
which remains unsupported and still exercises that code path.
Signed-off-by: Radhakrishnan P <gingeekrishna@gmail.com>

@RyanL1997RyanL1997 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM. Just would like to double check do we need to add doc test/doc for this?

@RyanL1997

Copy link
Copy Markdown
Collaborator

In addtion, I found out that

The following isn't something you introduced in this PR. The existing ROW_NUMBER case has the same shape, and ROW_NUMBER() OVER (ORDER BY age) already returns 1 for every row. The new RANK/DENSE_RANK cases inherit it.

Behaviour:

SELECT age, RANK() OVER (ORDER BY age) FROM employees
actual → 4, 4, 4, 4 (the partition size)
expected → 1, 2, 3, 4

Why: the new cases forward the caller's lowerBound/upperBound unchanged, and for SQL those default to WindowFrame.rowsUnbounded() = UNBOUNDED PRECEDING … UNBOUNDED FOLLOWING, so every row is ranked over the whole partition. Calcite's own SqlToRelConverter.convertOver forces UNBOUNDED PRECEDING … CURRENT ROW for operators with allowsFraming() == false; this branch needs to do the same.

One knock-on: the plan assertions can't catch it. allowsFraming() only suppresses the frame in the printed digest, not in the executed Window.Group — so RANK() OVER (ORDER BY $2 NULLS FIRST) prints identically whichever frame was built. Confirming this needs an execution-level assertion.

@dai-chen
dai-chen merged commit c4f27b2 into opensearch-project:mainAug 28, 2026
42 checks passed
gingeekrishna added a commit to gingeekrishna/sql that referenced this pull request Aug 30, 2026
WINDOW_FUNC_MAPPING (used by eventstats/streamstats) never supported
rank/dense_rank, but PPL's scalarWindowFunctionName grammar rule still
accepted the tokens, so `eventstats rank()` reached
CalciteRexNodeVisitor#visitWindowFunction and failed there with the
"not supported" message this PR improves. Meanwhile SQL's grammar
already accepts RANK()/DENSE_RANK() OVER (...), and opensearch-project#5720 is adding
real support for them on the SQL side via the same shared visitor.
Remove RANK/DENSE_RANK from scalarWindowFunctionName so PPL rejects
them at parse time instead of falling through to the shared
SQL/PPL visitor - this keeps the language separation at the parser
rather than relying on a WINDOW_FUNC_MAPPING check in shared planner
code, and avoids PPL silently gaining rank/dense_rank as a side effect
of opensearch-project#5720 registering them for SQL.
Updates the eventstats/streamstats tests added earlier in this PR to
expect a SyntaxCheckException (parse-time) instead of the semantic
"not supported" error, and switches the unrelated
visitWindowFunction-rejection unit test from rank() to percent_rank(),
which remains unsupported and still exercises that code path.
Signed-off-by: Radhakrishnan P <gingeekrishna@gmail.com>
gingeekrishna added a commit to gingeekrishna/sql that referenced this pull request Sep 2, 2026
…layer
opensearch-project#5720 added rank/dense_rank to the WINDOW_FUNC_MAPPING shared by both
SQL's RANK()/DENSE_RANK() OVER (...) and PPL's eventstats/streamstats,
which silently enabled them for eventstats/streamstats too:
testRankingWindowFunctionsUnsupportedInEventstats/InStreamstats
(added earlier in this PR) started failing with "expected
ResponseException to be thrown, but nothing was thrown" once opensearch-project#5720
merged, since eventstats rank() now builds a real window call instead
of hitting the "not supported" check.
That's a real gap, not just a test artifact: PPL's eventstats/
streamstats grammar has no ORDER BY syntax at all, so ranking has no
defined ordering to rank by there - unlike SQL's OVER(), which at
least has (optional) ORDER BY in its own clause.
A prior commit on this branch tried fixing this by removing RANK/
DENSE_RANK from the PPL grammar entirely, rejecting them at parse
time. That broke a different, cross-repo contract: OpenSearch-
Dashboards' PPL linter (validated by the "PPL grammar compatibility"
CI check) expects `eventstats rank()` to parse successfully and be
flagged by a semantic-layer diagnostic instead, so it could no longer
produce that diagnostic once the query stopped parsing. That commit
was reverted.
Fix this at the semantic layer instead, where it belongs: PPL only
ever reaches CalciteRexNodeVisitor#visitWindowFunction through
eventstats/streamstats (no other PPL syntax builds a WindowFunction
node), so context.queryType == PPL is an exact, unambiguous signal for
"this is an eventstats/streamstats call". Filter rank/dense_rank out
of the WINDOW_FUNC_MAPPING lookup specifically when queryType is PPL,
so they fall through to the existing "not supported in eventstats/
streamstats" error - exactly the pre-opensearch-project#5720 behavior - while leaving
SQL's RANK()/DENSE_RANK() OVER (...) handling (and everything else)
completely untouched.
Signed-off-by: Radhakrishnan P <gingeekrishna@gmail.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

analytic-engineenhancementNew feature or requestSQL

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@dai-chen@RyanL1997
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(sql): Support RANK/DENSE_RANK for unified SQL - #5720

Merged
dai-chen merged 1 commit into
opensearch-project:mainfrom
dai-chen:support-rank-dense-rank-calcite
Aug 28, 2026
Merged

feat(sql): Support RANK/DENSE_RANK for unified SQL#5720
dai-chen merged 1 commit into
opensearch-project:mainfrom
dai-chen:support-rank-dense-rank-calcite

Conversation

@dai-chen

@dai-chendai-chen commented Aug 24, 2026

Copy link
Copy Markdown
Collaborator

Description

RANK() and DENSE_RANK() were already accepted by the SQL grammar and declared as built-in function names, but were never registered as window functions, so planning failed with Unexpected window function: RANK. This PR registers both and maps them to Calcite's standard RANK and DENSE_RANK operators. No grammar, lexer or AST changes were needed.

Related Issues

Part of #5248

Check List

  • New functionality includes testing.
  • New functionality has been documented.
  • New functionality has javadoc added.
  • New functionality has a user manual doc added.
  • New PPL command checklist all confirmed.
  • API changes companion pull request created.
  • Commits are signed per the DCO using --signoff or -s.
  • Public documentation issue/PR created.

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.
For more information on following Developer Certificate of Origin and signing off your commits, please check here.

@dai-chendai-chen self-assigned this Aug 24, 2026
@github-actions

github-actionsBot commented Aug 24, 2026

Copy link
Copy Markdown
Contributor

PR Reviewer Guide 🔍

(Review updated until commit 544e467)

Here are some key observations to aid the review process:

🧪 PR contains tests
🔒 No security concerns identified
✅ No TODO sections
🔀 No multiple PR themes
⚡ No major issues detected

@github-actions

github-actionsBot commented Aug 24, 2026

Copy link
Copy Markdown
Contributor

PR Code Suggestions ✨

Latest suggestions up to 544e467

Explore these optional code suggestions:

CategorySuggestion Impact
General
Explicitly pass null for frame bounds

Since RANK and DENSE_RANK disallow framing, passing lowerBound and upperBound
parameters may cause unexpected behavior or errors. Consider passing null for both
bounds to explicitly indicate no framing is used.

core/src/main/java/org/opensearch/sql/calcite/utils/PlanUtils.java [236-251]

 case RANK:
return withOver(
context.relBuilder.aggregateCall(SqlStdOperatorTable.RANK),
partitions,
orderKeys,
false,
- lowerBound,- upperBound);+ null,+ null);
case DENSE_RANK:
return withOver(
context.relBuilder.aggregateCall(SqlStdOperatorTable.DENSE_RANK),
partitions,
orderKeys,
false,
- lowerBound,- upperBound);+ null,+ null);
Suggestion importance[1-10]: 7

__

Why: The suggestion correctly identifies that RANK and DENSE_RANK disallow framing, and the comment in the code confirms this. Passing null instead of lowerBound and upperBound would make the intent more explicit and prevent potential issues, though the current implementation may already handle this correctly through normalization.

Medium

Previous suggestions

Suggestions up to commit 77ae6e0
CategorySuggestion Impact
Possible issue
Explicitly pass null for frame bounds

Since RANK and DENSE_RANK disallow framing, passing lowerBound and upperBound
parameters may cause unexpected behavior or errors. Consider explicitly passing null
for both bounds to ensure proper handling of these ranking functions.

core/src/main/java/org/opensearch/sql/calcite/utils/PlanUtils.java [236-251]

 case RANK:
return withOver(
context.relBuilder.aggregateCall(SqlStdOperatorTable.RANK),
partitions,
orderKeys,
false,
- lowerBound,- upperBound);+ null,+ null);
case DENSE_RANK:
return withOver(
context.relBuilder.aggregateCall(SqlStdOperatorTable.DENSE_RANK),
partitions,
orderKeys,
false,
- lowerBound,- upperBound);+ null,+ null);
Suggestion importance[1-10]: 4

__

Why: While the suggestion correctly identifies that RANK and DENSE_RANK disallow framing, the comment in the PR already acknowledges this ("Calcite rank operators disallow framing, so the ROWS/RANGE flag below is normalized away"). The current implementation passes lowerBound and upperBound which are likely normalized by Calcite. Explicitly passing null would be slightly clearer but is not critical since the normalization handles this.

Low
Suggestions up to commit 747d1b6
CategorySuggestion Impact
Possible issue
Validate ORDER BY clause presence

The RANK and DENSE_RANK window functions require an ORDER BY clause to function
correctly. Consider validating that orderKeys is not empty before creating the
aggregate call to prevent runtime errors or unexpected behavior when no ordering is
specified.

core/src/main/java/org/opensearch/sql/calcite/utils/PlanUtils.java [236-251]

 case RANK:
+ if (orderKeys.isEmpty()) {+ throw new IllegalArgumentException("RANK requires ORDER BY clause");+ }
return withOver(
context.relBuilder.aggregateCall(SqlStdOperatorTable.RANK),
partitions,
orderKeys,
false,
lowerBound,
upperBound);
case DENSE_RANK:
+ if (orderKeys.isEmpty()) {+ throw new IllegalArgumentException("DENSE_RANK requires ORDER BY clause");+ }
return withOver(
context.relBuilder.aggregateCall(SqlStdOperatorTable.DENSE_RANK),
partitions,
orderKeys,
false,
lowerBound,
upperBound);
Suggestion importance[1-10]: 5

__

Why: While RANK and DENSE_RANK do require an ORDER BY clause semantically, this validation may already be handled at the SQL parsing/validation layer. Adding redundant validation here could be useful for defensive programming, but without evidence of missing validation upstream, this is a moderate improvement for robustness rather than fixing a critical bug.

Low
Suggestions up to commit 352ffbd
CategorySuggestion Impact
General
Use Set for multiple condition checks

Consider using a Set or EnumSet for checking multiple function names instead of
chained OR conditions. This improves readability and makes it easier to add more
functions in the future.

core/src/main/java/org/opensearch/sql/calcite/CalciteRexNodeVisitor.java [775-777]

-if (functionName == BuiltinFunctionName.ROW_NUMBER- || functionName == BuiltinFunctionName.RANK- || functionName == BuiltinFunctionName.DENSE_RANK) {+private static final Set<BuiltinFunctionName> NO_FIELD_WINDOW_FUNCTIONS = + EnumSet.of(BuiltinFunctionName.ROW_NUMBER, BuiltinFunctionName.RANK, BuiltinFunctionName.DENSE_RANK);+if (NO_FIELD_WINDOW_FUNCTIONS.contains(functionName)) {+
Suggestion importance[1-10]: 4

__

Why: While using an EnumSet would improve maintainability and readability, the current chained OR condition with three items is still acceptable and clear. The suggestion is valid but offers only a moderate improvement in code style rather than fixing a critical issue.

Low

@dai-chen
dai-chenforce-pushed the support-rank-dense-rank-calcite branch from 352ffbd to 747d1b6CompareAugust 25, 2026 00:21
@github-actions

Copy link
Copy Markdown
Contributor

Persistent review updated to latest commit 747d1b6

@dai-chen
dai-chenforce-pushed the support-rank-dense-rank-calcite branch from 747d1b6 to 77ae6e0CompareAugust 25, 2026 18:17
@github-actions

Copy link
Copy Markdown
Contributor

Persistent review updated to latest commit 77ae6e0

Both were accepted by the SQL grammar and declared as
BuiltinFunctionName constants, but missing from WINDOW_FUNC_MAPPING, so
visitWindowFunction rejected them with "Unexpected window function".
Register both, extend the ROW_NUMBER bypass of aggregate signature
validation since they likewise take no arguments, and lower them to
SqlStdOperatorTable.RANK and DENSE_RANK.
WINDOW_FUNC_MAPPING is shared, so PPL eventstats/streamstats now resolve
these functions as well.
Related to opensearch-project#5168
Signed-off-by: Chen Dai <daichen@amazon.com>
@dai-chen
dai-chenforce-pushed the support-rank-dense-rank-calcite branch from 77ae6e0 to 544e467CompareAugust 26, 2026 22:13
@github-actions

Copy link
Copy Markdown
Contributor

Persistent review updated to latest commit 544e467

@dai-chendai-chen changed the title feat(sql): Support RANK/DENSE_RANK in unified SQLfeat(sql): Support RANK/DENSE_RANK for unified SQLAug 26, 2026
@dai-chen
dai-chen marked this pull request as ready for review August 26, 2026 23:10
gingeekrishna added a commit to gingeekrishna/sql that referenced this pull request Aug 28, 2026
WINDOW_FUNC_MAPPING (used by eventstats/streamstats) never supported
rank/dense_rank, but PPL's scalarWindowFunctionName grammar rule still
accepted the tokens, so `eventstats rank()` reached
CalciteRexNodeVisitor#visitWindowFunction and failed there with the
"not supported" message this PR improves. Meanwhile SQL's grammar
already accepts RANK()/DENSE_RANK() OVER (...), and opensearch-project#5720 is adding
real support for them on the SQL side via the same shared visitor.
Remove RANK/DENSE_RANK from scalarWindowFunctionName so PPL rejects
them at parse time instead of falling through to the shared
SQL/PPL visitor - this keeps the language separation at the parser
rather than relying on a WINDOW_FUNC_MAPPING check in shared planner
code, and avoids PPL silently gaining rank/dense_rank as a side effect
of opensearch-project#5720 registering them for SQL.
Updates the eventstats/streamstats tests added earlier in this PR to
expect a SyntaxCheckException (parse-time) instead of the semantic
"not supported" error, and switches the unrelated
visitWindowFunction-rejection unit test from rank() to percent_rank(),
which remains unsupported and still exercises that code path.
Signed-off-by: Radhakrishnan P <gingeekrishna@gmail.com>

@RyanL1997RyanL1997 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM. Just would like to double check do we need to add doc test/doc for this?

@RyanL1997

Copy link
Copy Markdown
Collaborator

In addtion, I found out that

The following isn't something you introduced in this PR. The existing ROW_NUMBER case has the same shape, and ROW_NUMBER() OVER (ORDER BY age) already returns 1 for every row. The new RANK/DENSE_RANK cases inherit it.

Behaviour:

SELECT age, RANK() OVER (ORDER BY age) FROM employees
actual → 4, 4, 4, 4 (the partition size)
expected → 1, 2, 3, 4

Why: the new cases forward the caller's lowerBound/upperBound unchanged, and for SQL those default to WindowFrame.rowsUnbounded() = UNBOUNDED PRECEDING … UNBOUNDED FOLLOWING, so every row is ranked over the whole partition. Calcite's own SqlToRelConverter.convertOver forces UNBOUNDED PRECEDING … CURRENT ROW for operators with allowsFraming() == false; this branch needs to do the same.

One knock-on: the plan assertions can't catch it. allowsFraming() only suppresses the frame in the printed digest, not in the executed Window.Group — so RANK() OVER (ORDER BY $2 NULLS FIRST) prints identically whichever frame was built. Confirming this needs an execution-level assertion.

@dai-chen
dai-chen merged commit c4f27b2 into opensearch-project:mainAug 28, 2026
42 checks passed
gingeekrishna added a commit to gingeekrishna/sql that referenced this pull request Aug 30, 2026
WINDOW_FUNC_MAPPING (used by eventstats/streamstats) never supported
rank/dense_rank, but PPL's scalarWindowFunctionName grammar rule still
accepted the tokens, so `eventstats rank()` reached
CalciteRexNodeVisitor#visitWindowFunction and failed there with the
"not supported" message this PR improves. Meanwhile SQL's grammar
already accepts RANK()/DENSE_RANK() OVER (...), and opensearch-project#5720 is adding
real support for them on the SQL side via the same shared visitor.
Remove RANK/DENSE_RANK from scalarWindowFunctionName so PPL rejects
them at parse time instead of falling through to the shared
SQL/PPL visitor - this keeps the language separation at the parser
rather than relying on a WINDOW_FUNC_MAPPING check in shared planner
code, and avoids PPL silently gaining rank/dense_rank as a side effect
of opensearch-project#5720 registering them for SQL.
Updates the eventstats/streamstats tests added earlier in this PR to
expect a SyntaxCheckException (parse-time) instead of the semantic
"not supported" error, and switches the unrelated
visitWindowFunction-rejection unit test from rank() to percent_rank(),
which remains unsupported and still exercises that code path.
Signed-off-by: Radhakrishnan P <gingeekrishna@gmail.com>
gingeekrishna added a commit to gingeekrishna/sql that referenced this pull request Sep 2, 2026
…layer
opensearch-project#5720 added rank/dense_rank to the WINDOW_FUNC_MAPPING shared by both
SQL's RANK()/DENSE_RANK() OVER (...) and PPL's eventstats/streamstats,
which silently enabled them for eventstats/streamstats too:
testRankingWindowFunctionsUnsupportedInEventstats/InStreamstats
(added earlier in this PR) started failing with "expected
ResponseException to be thrown, but nothing was thrown" once opensearch-project#5720
merged, since eventstats rank() now builds a real window call instead
of hitting the "not supported" check.
That's a real gap, not just a test artifact: PPL's eventstats/
streamstats grammar has no ORDER BY syntax at all, so ranking has no
defined ordering to rank by there - unlike SQL's OVER(), which at
least has (optional) ORDER BY in its own clause.
A prior commit on this branch tried fixing this by removing RANK/
DENSE_RANK from the PPL grammar entirely, rejecting them at parse
time. That broke a different, cross-repo contract: OpenSearch-
Dashboards' PPL linter (validated by the "PPL grammar compatibility"
CI check) expects `eventstats rank()` to parse successfully and be
flagged by a semantic-layer diagnostic instead, so it could no longer
produce that diagnostic once the query stopped parsing. That commit
was reverted.
Fix this at the semantic layer instead, where it belongs: PPL only
ever reaches CalciteRexNodeVisitor#visitWindowFunction through
eventstats/streamstats (no other PPL syntax builds a WindowFunction
node), so context.queryType == PPL is an exact, unambiguous signal for
"this is an eventstats/streamstats call". Filter rank/dense_rank out
of the WINDOW_FUNC_MAPPING lookup specifically when queryType is PPL,
so they fall through to the existing "not supported in eventstats/
streamstats" error - exactly the pre-opensearch-project#5720 behavior - while leaving
SQL's RANK()/DENSE_RANK() OVER (...) handling (and everything else)
completely untouched.
Signed-off-by: Radhakrishnan P <gingeekrishna@gmail.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

analytic-engineenhancementNew feature or requestSQL

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@dai-chen@RyanL1997
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(sql): Support RANK/DENSE_RANK for unified SQL - #5720

Merged
dai-chen merged 1 commit into
opensearch-project:mainfrom
dai-chen:support-rank-dense-rank-calcite
Aug 28, 2026
Merged

feat(sql): Support RANK/DENSE_RANK for unified SQL#5720
dai-chen merged 1 commit into
opensearch-project:mainfrom
dai-chen:support-rank-dense-rank-calcite

Conversation

@dai-chen

@dai-chendai-chen commented Aug 24, 2026

Copy link
Copy Markdown
Collaborator

Description

RANK() and DENSE_RANK() were already accepted by the SQL grammar and declared as built-in function names, but were never registered as window functions, so planning failed with Unexpected window function: RANK. This PR registers both and maps them to Calcite's standard RANK and DENSE_RANK operators. No grammar, lexer or AST changes were needed.

Related Issues

Part of #5248

Check List

  • New functionality includes testing.
  • New functionality has been documented.
  • New functionality has javadoc added.
  • New functionality has a user manual doc added.
  • New PPL command checklist all confirmed.
  • API changes companion pull request created.
  • Commits are signed per the DCO using --signoff or -s.
  • Public documentation issue/PR created.

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.
For more information on following Developer Certificate of Origin and signing off your commits, please check here.

@dai-chendai-chen self-assigned this Aug 24, 2026
@github-actions

github-actionsBot commented Aug 24, 2026

Copy link
Copy Markdown
Contributor

PR Reviewer Guide 🔍

(Review updated until commit 544e467)

Here are some key observations to aid the review process:

🧪 PR contains tests
🔒 No security concerns identified
✅ No TODO sections
🔀 No multiple PR themes
⚡ No major issues detected

@github-actions

github-actionsBot commented Aug 24, 2026

Copy link
Copy Markdown
Contributor

PR Code Suggestions ✨

Latest suggestions up to 544e467

Explore these optional code suggestions:

CategorySuggestion Impact
General
Explicitly pass null for frame bounds

Since RANK and DENSE_RANK disallow framing, passing lowerBound and upperBound
parameters may cause unexpected behavior or errors. Consider passing null for both
bounds to explicitly indicate no framing is used.

core/src/main/java/org/opensearch/sql/calcite/utils/PlanUtils.java [236-251]

 case RANK:
return withOver(
context.relBuilder.aggregateCall(SqlStdOperatorTable.RANK),
partitions,
orderKeys,
false,
- lowerBound,- upperBound);+ null,+ null);
case DENSE_RANK:
return withOver(
context.relBuilder.aggregateCall(SqlStdOperatorTable.DENSE_RANK),
partitions,
orderKeys,
false,
- lowerBound,- upperBound);+ null,+ null);
Suggestion importance[1-10]: 7

__

Why: The suggestion correctly identifies that RANK and DENSE_RANK disallow framing, and the comment in the code confirms this. Passing null instead of lowerBound and upperBound would make the intent more explicit and prevent potential issues, though the current implementation may already handle this correctly through normalization.

Medium

Previous suggestions

Suggestions up to commit 77ae6e0
CategorySuggestion Impact
Possible issue
Explicitly pass null for frame bounds

Since RANK and DENSE_RANK disallow framing, passing lowerBound and upperBound
parameters may cause unexpected behavior or errors. Consider explicitly passing null
for both bounds to ensure proper handling of these ranking functions.

core/src/main/java/org/opensearch/sql/calcite/utils/PlanUtils.java [236-251]

 case RANK:
return withOver(
context.relBuilder.aggregateCall(SqlStdOperatorTable.RANK),
partitions,
orderKeys,
false,
- lowerBound,- upperBound);+ null,+ null);
case DENSE_RANK:
return withOver(
context.relBuilder.aggregateCall(SqlStdOperatorTable.DENSE_RANK),
partitions,
orderKeys,
false,
- lowerBound,- upperBound);+ null,+ null);
Suggestion importance[1-10]: 4

__

Why: While the suggestion correctly identifies that RANK and DENSE_RANK disallow framing, the comment in the PR already acknowledges this ("Calcite rank operators disallow framing, so the ROWS/RANGE flag below is normalized away"). The current implementation passes lowerBound and upperBound which are likely normalized by Calcite. Explicitly passing null would be slightly clearer but is not critical since the normalization handles this.

Low
Suggestions up to commit 747d1b6
CategorySuggestion Impact
Possible issue
Validate ORDER BY clause presence

The RANK and DENSE_RANK window functions require an ORDER BY clause to function
correctly. Consider validating that orderKeys is not empty before creating the
aggregate call to prevent runtime errors or unexpected behavior when no ordering is
specified.

core/src/main/java/org/opensearch/sql/calcite/utils/PlanUtils.java [236-251]

 case RANK:
+ if (orderKeys.isEmpty()) {+ throw new IllegalArgumentException("RANK requires ORDER BY clause");+ }
return withOver(
context.relBuilder.aggregateCall(SqlStdOperatorTable.RANK),
partitions,
orderKeys,
false,
lowerBound,
upperBound);
case DENSE_RANK:
+ if (orderKeys.isEmpty()) {+ throw new IllegalArgumentException("DENSE_RANK requires ORDER BY clause");+ }
return withOver(
context.relBuilder.aggregateCall(SqlStdOperatorTable.DENSE_RANK),
partitions,
orderKeys,
false,
lowerBound,
upperBound);
Suggestion importance[1-10]: 5

__

Why: While RANK and DENSE_RANK do require an ORDER BY clause semantically, this validation may already be handled at the SQL parsing/validation layer. Adding redundant validation here could be useful for defensive programming, but without evidence of missing validation upstream, this is a moderate improvement for robustness rather than fixing a critical bug.

Low
Suggestions up to commit 352ffbd
CategorySuggestion Impact
General
Use Set for multiple condition checks

Consider using a Set or EnumSet for checking multiple function names instead of
chained OR conditions. This improves readability and makes it easier to add more
functions in the future.

core/src/main/java/org/opensearch/sql/calcite/CalciteRexNodeVisitor.java [775-777]

-if (functionName == BuiltinFunctionName.ROW_NUMBER- || functionName == BuiltinFunctionName.RANK- || functionName == BuiltinFunctionName.DENSE_RANK) {+private static final Set<BuiltinFunctionName> NO_FIELD_WINDOW_FUNCTIONS = + EnumSet.of(BuiltinFunctionName.ROW_NUMBER, BuiltinFunctionName.RANK, BuiltinFunctionName.DENSE_RANK);+if (NO_FIELD_WINDOW_FUNCTIONS.contains(functionName)) {+
Suggestion importance[1-10]: 4

__

Why: While using an EnumSet would improve maintainability and readability, the current chained OR condition with three items is still acceptable and clear. The suggestion is valid but offers only a moderate improvement in code style rather than fixing a critical issue.

Low

@dai-chen
dai-chenforce-pushed the support-rank-dense-rank-calcite branch from 352ffbd to 747d1b6CompareAugust 25, 2026 00:21
@github-actions

Copy link
Copy Markdown
Contributor

Persistent review updated to latest commit 747d1b6

@dai-chen
dai-chenforce-pushed the support-rank-dense-rank-calcite branch from 747d1b6 to 77ae6e0CompareAugust 25, 2026 18:17
@github-actions

Copy link
Copy Markdown
Contributor

Persistent review updated to latest commit 77ae6e0

Both were accepted by the SQL grammar and declared as
BuiltinFunctionName constants, but missing from WINDOW_FUNC_MAPPING, so
visitWindowFunction rejected them with "Unexpected window function".
Register both, extend the ROW_NUMBER bypass of aggregate signature
validation since they likewise take no arguments, and lower them to
SqlStdOperatorTable.RANK and DENSE_RANK.
WINDOW_FUNC_MAPPING is shared, so PPL eventstats/streamstats now resolve
these functions as well.
Related to opensearch-project#5168
Signed-off-by: Chen Dai <daichen@amazon.com>
@dai-chen
dai-chenforce-pushed the support-rank-dense-rank-calcite branch from 77ae6e0 to 544e467CompareAugust 26, 2026 22:13
@github-actions

Copy link
Copy Markdown
Contributor

Persistent review updated to latest commit 544e467

@dai-chendai-chen changed the title feat(sql): Support RANK/DENSE_RANK in unified SQLfeat(sql): Support RANK/DENSE_RANK for unified SQLAug 26, 2026
@dai-chen
dai-chen marked this pull request as ready for review August 26, 2026 23:10
gingeekrishna added a commit to gingeekrishna/sql that referenced this pull request Aug 28, 2026
WINDOW_FUNC_MAPPING (used by eventstats/streamstats) never supported
rank/dense_rank, but PPL's scalarWindowFunctionName grammar rule still
accepted the tokens, so `eventstats rank()` reached
CalciteRexNodeVisitor#visitWindowFunction and failed there with the
"not supported" message this PR improves. Meanwhile SQL's grammar
already accepts RANK()/DENSE_RANK() OVER (...), and opensearch-project#5720 is adding
real support for them on the SQL side via the same shared visitor.
Remove RANK/DENSE_RANK from scalarWindowFunctionName so PPL rejects
them at parse time instead of falling through to the shared
SQL/PPL visitor - this keeps the language separation at the parser
rather than relying on a WINDOW_FUNC_MAPPING check in shared planner
code, and avoids PPL silently gaining rank/dense_rank as a side effect
of opensearch-project#5720 registering them for SQL.
Updates the eventstats/streamstats tests added earlier in this PR to
expect a SyntaxCheckException (parse-time) instead of the semantic
"not supported" error, and switches the unrelated
visitWindowFunction-rejection unit test from rank() to percent_rank(),
which remains unsupported and still exercises that code path.
Signed-off-by: Radhakrishnan P <gingeekrishna@gmail.com>

@RyanL1997RyanL1997 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM. Just would like to double check do we need to add doc test/doc for this?

@RyanL1997

Copy link
Copy Markdown
Collaborator

In addtion, I found out that

The following isn't something you introduced in this PR. The existing ROW_NUMBER case has the same shape, and ROW_NUMBER() OVER (ORDER BY age) already returns 1 for every row. The new RANK/DENSE_RANK cases inherit it.

Behaviour:

SELECT age, RANK() OVER (ORDER BY age) FROM employees
actual → 4, 4, 4, 4 (the partition size)
expected → 1, 2, 3, 4

Why: the new cases forward the caller's lowerBound/upperBound unchanged, and for SQL those default to WindowFrame.rowsUnbounded() = UNBOUNDED PRECEDING … UNBOUNDED FOLLOWING, so every row is ranked over the whole partition. Calcite's own SqlToRelConverter.convertOver forces UNBOUNDED PRECEDING … CURRENT ROW for operators with allowsFraming() == false; this branch needs to do the same.

One knock-on: the plan assertions can't catch it. allowsFraming() only suppresses the frame in the printed digest, not in the executed Window.Group — so RANK() OVER (ORDER BY $2 NULLS FIRST) prints identically whichever frame was built. Confirming this needs an execution-level assertion.

@dai-chen
dai-chen merged commit c4f27b2 into opensearch-project:mainAug 28, 2026
42 checks passed
gingeekrishna added a commit to gingeekrishna/sql that referenced this pull request Aug 30, 2026
WINDOW_FUNC_MAPPING (used by eventstats/streamstats) never supported
rank/dense_rank, but PPL's scalarWindowFunctionName grammar rule still
accepted the tokens, so `eventstats rank()` reached
CalciteRexNodeVisitor#visitWindowFunction and failed there with the
"not supported" message this PR improves. Meanwhile SQL's grammar
already accepts RANK()/DENSE_RANK() OVER (...), and opensearch-project#5720 is adding
real support for them on the SQL side via the same shared visitor.
Remove RANK/DENSE_RANK from scalarWindowFunctionName so PPL rejects
them at parse time instead of falling through to the shared
SQL/PPL visitor - this keeps the language separation at the parser
rather than relying on a WINDOW_FUNC_MAPPING check in shared planner
code, and avoids PPL silently gaining rank/dense_rank as a side effect
of opensearch-project#5720 registering them for SQL.
Updates the eventstats/streamstats tests added earlier in this PR to
expect a SyntaxCheckException (parse-time) instead of the semantic
"not supported" error, and switches the unrelated
visitWindowFunction-rejection unit test from rank() to percent_rank(),
which remains unsupported and still exercises that code path.
Signed-off-by: Radhakrishnan P <gingeekrishna@gmail.com>
gingeekrishna added a commit to gingeekrishna/sql that referenced this pull request Sep 2, 2026
…layer
opensearch-project#5720 added rank/dense_rank to the WINDOW_FUNC_MAPPING shared by both
SQL's RANK()/DENSE_RANK() OVER (...) and PPL's eventstats/streamstats,
which silently enabled them for eventstats/streamstats too:
testRankingWindowFunctionsUnsupportedInEventstats/InStreamstats
(added earlier in this PR) started failing with "expected
ResponseException to be thrown, but nothing was thrown" once opensearch-project#5720
merged, since eventstats rank() now builds a real window call instead
of hitting the "not supported" check.
That's a real gap, not just a test artifact: PPL's eventstats/
streamstats grammar has no ORDER BY syntax at all, so ranking has no
defined ordering to rank by there - unlike SQL's OVER(), which at
least has (optional) ORDER BY in its own clause.
A prior commit on this branch tried fixing this by removing RANK/
DENSE_RANK from the PPL grammar entirely, rejecting them at parse
time. That broke a different, cross-repo contract: OpenSearch-
Dashboards' PPL linter (validated by the "PPL grammar compatibility"
CI check) expects `eventstats rank()` to parse successfully and be
flagged by a semantic-layer diagnostic instead, so it could no longer
produce that diagnostic once the query stopped parsing. That commit
was reverted.
Fix this at the semantic layer instead, where it belongs: PPL only
ever reaches CalciteRexNodeVisitor#visitWindowFunction through
eventstats/streamstats (no other PPL syntax builds a WindowFunction
node), so context.queryType == PPL is an exact, unambiguous signal for
"this is an eventstats/streamstats call". Filter rank/dense_rank out
of the WINDOW_FUNC_MAPPING lookup specifically when queryType is PPL,
so they fall through to the existing "not supported in eventstats/
streamstats" error - exactly the pre-opensearch-project#5720 behavior - while leaving
SQL's RANK()/DENSE_RANK() OVER (...) handling (and everything else)
completely untouched.
Signed-off-by: Radhakrishnan P <gingeekrishna@gmail.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

analytic-engineenhancementNew feature or requestSQL

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@dai-chen@RyanL1997
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

feat(sql): Support RANK/DENSE_RANK for unified SQL - #5720

Merged
dai-chen merged 1 commit into
opensearch-project:mainfrom
dai-chen:support-rank-dense-rank-calcite
Aug 28, 2026
Merged

feat(sql): Support RANK/DENSE_RANK for unified SQL#5720
dai-chen merged 1 commit into
opensearch-project:mainfrom
dai-chen:support-rank-dense-rank-calcite

Conversation

@dai-chen

@dai-chendai-chen commented Aug 24, 2026

Copy link
Copy Markdown
Collaborator

Description

RANK() and DENSE_RANK() were already accepted by the SQL grammar and declared as built-in function names, but were never registered as window functions, so planning failed with Unexpected window function: RANK. This PR registers both and maps them to Calcite's standard RANK and DENSE_RANK operators. No grammar, lexer or AST changes were needed.

Related Issues

Part of #5248

Check List

  • New functionality includes testing.
  • New functionality has been documented.
  • New functionality has javadoc added.
  • New functionality has a user manual doc added.
  • New PPL command checklist all confirmed.
  • API changes companion pull request created.
  • Commits are signed per the DCO using --signoff or -s.
  • Public documentation issue/PR created.

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.
For more information on following Developer Certificate of Origin and signing off your commits, please check here.

@dai-chendai-chen self-assigned this Aug 24, 2026
@github-actions

github-actionsBot commented Aug 24, 2026

Copy link
Copy Markdown
Contributor

PR Reviewer Guide 🔍

(Review updated until commit 544e467)

Here are some key observations to aid the review process:

🧪 PR contains tests
🔒 No security concerns identified
✅ No TODO sections
🔀 No multiple PR themes
⚡ No major issues detected

@github-actions

github-actionsBot commented Aug 24, 2026

Copy link
Copy Markdown
Contributor

PR Code Suggestions ✨

Latest suggestions up to 544e467

Explore these optional code suggestions:

CategorySuggestion Impact
General
Explicitly pass null for frame bounds

Since RANK and DENSE_RANK disallow framing, passing lowerBound and upperBound
parameters may cause unexpected behavior or errors. Consider passing null for both
bounds to explicitly indicate no framing is used.

core/src/main/java/org/opensearch/sql/calcite/utils/PlanUtils.java [236-251]

 case RANK:
return withOver(
context.relBuilder.aggregateCall(SqlStdOperatorTable.RANK),
partitions,
orderKeys,
false,
- lowerBound,- upperBound);+ null,+ null);
case DENSE_RANK:
return withOver(
context.relBuilder.aggregateCall(SqlStdOperatorTable.DENSE_RANK),
partitions,
orderKeys,
false,
- lowerBound,- upperBound);+ null,+ null);
Suggestion importance[1-10]: 7

__

Why: The suggestion correctly identifies that RANK and DENSE_RANK disallow framing, and the comment in the code confirms this. Passing null instead of lowerBound and upperBound would make the intent more explicit and prevent potential issues, though the current implementation may already handle this correctly through normalization.

Medium

Previous suggestions

Suggestions up to commit 77ae6e0
CategorySuggestion Impact
Possible issue
Explicitly pass null for frame bounds

Since RANK and DENSE_RANK disallow framing, passing lowerBound and upperBound
parameters may cause unexpected behavior or errors. Consider explicitly passing null
for both bounds to ensure proper handling of these ranking functions.

core/src/main/java/org/opensearch/sql/calcite/utils/PlanUtils.java [236-251]

 case RANK:
return withOver(
context.relBuilder.aggregateCall(SqlStdOperatorTable.RANK),
partitions,
orderKeys,
false,
- lowerBound,- upperBound);+ null,+ null);
case DENSE_RANK:
return withOver(
context.relBuilder.aggregateCall(SqlStdOperatorTable.DENSE_RANK),
partitions,
orderKeys,
false,
- lowerBound,- upperBound);+ null,+ null);
Suggestion importance[1-10]: 4

__

Why: While the suggestion correctly identifies that RANK and DENSE_RANK disallow framing, the comment in the PR already acknowledges this ("Calcite rank operators disallow framing, so the ROWS/RANGE flag below is normalized away"). The current implementation passes lowerBound and upperBound which are likely normalized by Calcite. Explicitly passing null would be slightly clearer but is not critical since the normalization handles this.

Low
Suggestions up to commit 747d1b6
CategorySuggestion Impact
Possible issue
Validate ORDER BY clause presence

The RANK and DENSE_RANK window functions require an ORDER BY clause to function
correctly. Consider validating that orderKeys is not empty before creating the
aggregate call to prevent runtime errors or unexpected behavior when no ordering is
specified.

core/src/main/java/org/opensearch/sql/calcite/utils/PlanUtils.java [236-251]

 case RANK:
+ if (orderKeys.isEmpty()) {+ throw new IllegalArgumentException("RANK requires ORDER BY clause");+ }
return withOver(
context.relBuilder.aggregateCall(SqlStdOperatorTable.RANK),
partitions,
orderKeys,
false,
lowerBound,
upperBound);
case DENSE_RANK:
+ if (orderKeys.isEmpty()) {+ throw new IllegalArgumentException("DENSE_RANK requires ORDER BY clause");+ }
return withOver(
context.relBuilder.aggregateCall(SqlStdOperatorTable.DENSE_RANK),
partitions,
orderKeys,
false,
lowerBound,
upperBound);
Suggestion importance[1-10]: 5

__

Why: While RANK and DENSE_RANK do require an ORDER BY clause semantically, this validation may already be handled at the SQL parsing/validation layer. Adding redundant validation here could be useful for defensive programming, but without evidence of missing validation upstream, this is a moderate improvement for robustness rather than fixing a critical bug.

Low
Suggestions up to commit 352ffbd
CategorySuggestion Impact
General
Use Set for multiple condition checks

Consider using a Set or EnumSet for checking multiple function names instead of
chained OR conditions. This improves readability and makes it easier to add more
functions in the future.

core/src/main/java/org/opensearch/sql/calcite/CalciteRexNodeVisitor.java [775-777]

-if (functionName == BuiltinFunctionName.ROW_NUMBER- || functionName == BuiltinFunctionName.RANK- || functionName == BuiltinFunctionName.DENSE_RANK) {+private static final Set<BuiltinFunctionName> NO_FIELD_WINDOW_FUNCTIONS = + EnumSet.of(BuiltinFunctionName.ROW_NUMBER, BuiltinFunctionName.RANK, BuiltinFunctionName.DENSE_RANK);+if (NO_FIELD_WINDOW_FUNCTIONS.contains(functionName)) {+
Suggestion importance[1-10]: 4

__

Why: While using an EnumSet would improve maintainability and readability, the current chained OR condition with three items is still acceptable and clear. The suggestion is valid but offers only a moderate improvement in code style rather than fixing a critical issue.

Low

@dai-chen
dai-chenforce-pushed the support-rank-dense-rank-calcite branch from 352ffbd to 747d1b6CompareAugust 25, 2026 00:21
@github-actions

Copy link
Copy Markdown
Contributor

Persistent review updated to latest commit 747d1b6

@dai-chen
dai-chenforce-pushed the support-rank-dense-rank-calcite branch from 747d1b6 to 77ae6e0CompareAugust 25, 2026 18:17
@github-actions

Copy link
Copy Markdown
Contributor

Persistent review updated to latest commit 77ae6e0

Both were accepted by the SQL grammar and declared as
BuiltinFunctionName constants, but missing from WINDOW_FUNC_MAPPING, so
visitWindowFunction rejected them with "Unexpected window function".
Register both, extend the ROW_NUMBER bypass of aggregate signature
validation since they likewise take no arguments, and lower them to
SqlStdOperatorTable.RANK and DENSE_RANK.
WINDOW_FUNC_MAPPING is shared, so PPL eventstats/streamstats now resolve
these functions as well.
Related to opensearch-project#5168
Signed-off-by: Chen Dai <daichen@amazon.com>
@dai-chen
dai-chenforce-pushed the support-rank-dense-rank-calcite branch from 77ae6e0 to 544e467CompareAugust 26, 2026 22:13
@github-actions

Copy link
Copy Markdown
Contributor

Persistent review updated to latest commit 544e467

@dai-chendai-chen changed the title feat(sql): Support RANK/DENSE_RANK in unified SQLfeat(sql): Support RANK/DENSE_RANK for unified SQLAug 26, 2026
@dai-chen
dai-chen marked this pull request as ready for review August 26, 2026 23:10
gingeekrishna added a commit to gingeekrishna/sql that referenced this pull request Aug 28, 2026
WINDOW_FUNC_MAPPING (used by eventstats/streamstats) never supported
rank/dense_rank, but PPL's scalarWindowFunctionName grammar rule still
accepted the tokens, so `eventstats rank()` reached
CalciteRexNodeVisitor#visitWindowFunction and failed there with the
"not supported" message this PR improves. Meanwhile SQL's grammar
already accepts RANK()/DENSE_RANK() OVER (...), and opensearch-project#5720 is adding
real support for them on the SQL side via the same shared visitor.
Remove RANK/DENSE_RANK from scalarWindowFunctionName so PPL rejects
them at parse time instead of falling through to the shared
SQL/PPL visitor - this keeps the language separation at the parser
rather than relying on a WINDOW_FUNC_MAPPING check in shared planner
code, and avoids PPL silently gaining rank/dense_rank as a side effect
of opensearch-project#5720 registering them for SQL.
Updates the eventstats/streamstats tests added earlier in this PR to
expect a SyntaxCheckException (parse-time) instead of the semantic
"not supported" error, and switches the unrelated
visitWindowFunction-rejection unit test from rank() to percent_rank(),
which remains unsupported and still exercises that code path.
Signed-off-by: Radhakrishnan P <gingeekrishna@gmail.com>

@RyanL1997RyanL1997 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM. Just would like to double check do we need to add doc test/doc for this?

@RyanL1997

Copy link
Copy Markdown
Collaborator

In addtion, I found out that

The following isn't something you introduced in this PR. The existing ROW_NUMBER case has the same shape, and ROW_NUMBER() OVER (ORDER BY age) already returns 1 for every row. The new RANK/DENSE_RANK cases inherit it.

Behaviour:

SELECT age, RANK() OVER (ORDER BY age) FROM employees
actual → 4, 4, 4, 4 (the partition size)
expected → 1, 2, 3, 4

Why: the new cases forward the caller's lowerBound/upperBound unchanged, and for SQL those default to WindowFrame.rowsUnbounded() = UNBOUNDED PRECEDING … UNBOUNDED FOLLOWING, so every row is ranked over the whole partition. Calcite's own SqlToRelConverter.convertOver forces UNBOUNDED PRECEDING … CURRENT ROW for operators with allowsFraming() == false; this branch needs to do the same.

One knock-on: the plan assertions can't catch it. allowsFraming() only suppresses the frame in the printed digest, not in the executed Window.Group — so RANK() OVER (ORDER BY $2 NULLS FIRST) prints identically whichever frame was built. Confirming this needs an execution-level assertion.

@dai-chen
dai-chen merged commit c4f27b2 into opensearch-project:mainAug 28, 2026
42 checks passed
gingeekrishna added a commit to gingeekrishna/sql that referenced this pull request Aug 30, 2026
WINDOW_FUNC_MAPPING (used by eventstats/streamstats) never supported
rank/dense_rank, but PPL's scalarWindowFunctionName grammar rule still
accepted the tokens, so `eventstats rank()` reached
CalciteRexNodeVisitor#visitWindowFunction and failed there with the
"not supported" message this PR improves. Meanwhile SQL's grammar
already accepts RANK()/DENSE_RANK() OVER (...), and opensearch-project#5720 is adding
real support for them on the SQL side via the same shared visitor.
Remove RANK/DENSE_RANK from scalarWindowFunctionName so PPL rejects
them at parse time instead of falling through to the shared
SQL/PPL visitor - this keeps the language separation at the parser
rather than relying on a WINDOW_FUNC_MAPPING check in shared planner
code, and avoids PPL silently gaining rank/dense_rank as a side effect
of opensearch-project#5720 registering them for SQL.
Updates the eventstats/streamstats tests added earlier in this PR to
expect a SyntaxCheckException (parse-time) instead of the semantic
"not supported" error, and switches the unrelated
visitWindowFunction-rejection unit test from rank() to percent_rank(),
which remains unsupported and still exercises that code path.
Signed-off-by: Radhakrishnan P <gingeekrishna@gmail.com>
gingeekrishna added a commit to gingeekrishna/sql that referenced this pull request Sep 2, 2026
…layer
opensearch-project#5720 added rank/dense_rank to the WINDOW_FUNC_MAPPING shared by both
SQL's RANK()/DENSE_RANK() OVER (...) and PPL's eventstats/streamstats,
which silently enabled them for eventstats/streamstats too:
testRankingWindowFunctionsUnsupportedInEventstats/InStreamstats
(added earlier in this PR) started failing with "expected
ResponseException to be thrown, but nothing was thrown" once opensearch-project#5720
merged, since eventstats rank() now builds a real window call instead
of hitting the "not supported" check.
That's a real gap, not just a test artifact: PPL's eventstats/
streamstats grammar has no ORDER BY syntax at all, so ranking has no
defined ordering to rank by there - unlike SQL's OVER(), which at
least has (optional) ORDER BY in its own clause.
A prior commit on this branch tried fixing this by removing RANK/
DENSE_RANK from the PPL grammar entirely, rejecting them at parse
time. That broke a different, cross-repo contract: OpenSearch-
Dashboards' PPL linter (validated by the "PPL grammar compatibility"
CI check) expects `eventstats rank()` to parse successfully and be
flagged by a semantic-layer diagnostic instead, so it could no longer
produce that diagnostic once the query stopped parsing. That commit
was reverted.
Fix this at the semantic layer instead, where it belongs: PPL only
ever reaches CalciteRexNodeVisitor#visitWindowFunction through
eventstats/streamstats (no other PPL syntax builds a WindowFunction
node), so context.queryType == PPL is an exact, unambiguous signal for
"this is an eventstats/streamstats call". Filter rank/dense_rank out
of the WINDOW_FUNC_MAPPING lookup specifically when queryType is PPL,
so they fall through to the existing "not supported in eventstats/
streamstats" error - exactly the pre-opensearch-project#5720 behavior - while leaving
SQL's RANK()/DENSE_RANK() OVER (...) handling (and everything else)
completely untouched.
Signed-off-by: Radhakrishnan P <gingeekrishna@gmail.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

analytic-engineenhancementNew feature or requestSQL

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@dai-chen@RyanL1997