Use _doc + _shard_doc as sort tiebreaker to get better performance - #4569

Merged
penghuo merged 2 commits into
opensearch-project:mainfrom
LantaoJin:pr/sort_tiebreaker
Oct 15, 2025
Merged

Use _doc + _shard_doc as sort tiebreaker to get better performance#4569
penghuo merged 2 commits into
opensearch-project:mainfrom
LantaoJin:pr/sort_tiebreaker

Conversation

@LantaoJin

@LantaoJinLantaoJin commented Oct 15, 2025

Copy link
Copy Markdown
Member

Description

Before #4378, the sort in PIT search is
case 1: if no sort field specified, sort by _doc + _id (+ means "then"). (❎ could cause high memory issue)
case 2: if sort fields specified, sort by fields. (❎ paged results could miss or duplicate hits)
case 3: if sort fields specified and query contains a filter, sort by _doc. (❎ paged results could miss or duplicate hits)

#4378 added the _shard_doc as sort tiebreaker with
case 1: if no sort field specified, sort by _shard_doc. (❎ performance regression)
case 2: if sort fields specified, sort by fields + _shard_doc.(❎ lower performance on low cardinality field)

#4435 found performance regression in case 1 and partially revert the changes to
case 1: if no sort field specified, sort by _doc + _id. (❎ could cause high memory issue)
case 2: if sort fields specified, sort by fields. (❎ paged results could miss or duplicate hits)

After this PR, we change the sort in PIT search to
case 1: if no sort field specified, sort by _doc + _shard_doc. ✅
case 2: if sort fields specified, sort by fields + _doc + _shard_doc.✅

RCA of performance regression:
_shard_doc is not a stored field in index which will be generated in runtime when comparison. Computing _shard_doc per document is a high cost operation. But sorting by _doc then _shard_doc only generates _shard_doc when the _doc values are conflicted.
Even in the case of user specified sort fields, we should sort by fields then _doc then _shard_doc to reduce the computing of _shard_doc. For example, if the sort field is a low cardinality field, e.g. gender, sorting by gender then _doc then _shard_doc generates _shard_doc for comparison only if values of gender and _doc are both conflicted.

This PR is no needed to backport to 2.19-dev since shard_doc feature is only available since OS 3.3.0

Related Issues

Resolves #[Issue number to be closed when this PR is merged]

Check List

  • New functionality includes testing.
  • New functionality has been documented.
  • New functionality has javadoc added.
  • New functionality has a user manual doc added.
  • New PPL command checklist all confirmed.
  • API changes companion pull request created.
  • Commits are signed per the DCO using --signoff or -s.
  • Public documentation issue/PR created.

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.
For more information on following Developer Certificate of Origin and signing off your commits, please check here.

Signed-off-by: Lantao Jin <ltjin@amazon.com>
@LantaoJinLantaoJin added the enhancement New feature or request label Oct 15, 2025
@LantaoJinLantaoJin changed the title Use _shard_doc as sort tiebreaker to get better performanceUse _doc + _shard_doc as sort tiebreaker to get better performanceOct 15, 2025
Signed-off-by: Lantao Jin <ltjin@amazon.com>
// Workaround to preserve sort location more exactly,
// see https://github.com/opensearch-project/sql/pull/3061
this.sourceBuilder.sort(METADATA_FIELD_ID, ASC);
this.sourceBuilder.sort(SortBuilders.shardDocSort());

@SwiddisSwiddisOct 15, 2025

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Does it matter if we duplicate fields in the sorting list? We could simplify/remove the below else logic by just always appending this, I would expect Lucene to optimize it in the background but I haven't measured it.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

not sure will Lucene optimize duplicated fields or _doc in sorting, but for sure the duplicated _shard_doc is not allowed in OpenSearch Core. It is no harmful for restricted checker here.

@anasalkouz

Copy link
Copy Markdown
Member

Can you share the performance benchmark for the 3 approaches?

@LantaoJin

LantaoJin commented Oct 15, 2025

Copy link
Copy Markdown
MemberAuthor

Can you share the performance benchmark for the 3 approaches?

I haven't run the benchmark, the RCA was made by reading the code of Luence and OS shard_doc feature.

The performance of _doc then _shard_doc is same as _doc then _id, provided by @ahkcs on Oct 1st. (the case 1)

For case 2, the current fields + _doc + _shard_doc is an further optimization upon fields + _shard_doc based on above benchmark result with inference.

Will rerun some benchmark to double confirm.

@penghuo
penghuo merged commit 3388dc7 into opensearch-project:mainOct 15, 2025
38 checks passed
ykmr1224 added a commit to ykmr1224/sql that referenced this pull request Oct 15, 2025
commit cba8d02
Author: Tomoyuki MORITA <moritato@amazon.com>
Date: Wed Oct 15 13:08:05 2025 -0700
Add MAP_APPEND internal function to Calcite PPL (opensearch-project#4515)
* Add MAP_APPEND internal function to Calcite PPL
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
* Minor fix
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
* Address comment
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
* Rebase and fix IT issue
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
---------
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
commit 3388dc7
Author: Lantao Jin <ltjin@amazon.com>
Date: Thu Oct 16 01:45:29 2025 +0800
Use `_doc` + `_shard_doc` as sort tiebreaker to get better performance (opensearch-project#4569)
* Use _shard_doc as sort tiebreaker
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* _doc as a part of tie-breaker have better performance
Signed-off-by: Lantao Jin <ltjin@amazon.com>
---------
Signed-off-by: Lantao Jin <ltjin@amazon.com>
commit 5630119
Author: qianheng <qianheng@amazon.com>
Date: Wed Oct 15 16:40:41 2025 +0800
Fix sort push down into agg after project already pushed (opensearch-project#4546)
* Fix sort push down into agg
Signed-off-by: Heng Qian <qianheng@amazon.com>
* Change some json files to yaml format
Signed-off-by: Heng Qian <qianheng@amazon.com>
---------
Signed-off-by: Heng Qian <qianheng@amazon.com>
commit 1e62fba
Author: Tomoyuki MORITA <moritato@amazon.com>
Date: Tue Oct 14 17:20:38 2025 -0700
Fix JsonExtractAllFunctionIT failure (opensearch-project#4556)
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
commit 02ee33e
Author: Kai Huang <105710027+ahkcs@users.noreply.github.com>
Date: Tue Oct 14 14:28:53 2025 -0700
Add more examples to the `where` command doc (opensearch-project#4457)
Co-authored-by: Manasvini B S <manasvis@amazon.com>
commit 0b7e86c
Author: Jialiang Liang <jiallian@amazon.com>
Date: Tue Oct 14 10:46:01 2025 -0700
[Enhancement] Error handling for illegal character usage in java regex named capture group (opensearch-project#4434)
Co-authored-by: Simeon Widdis <sawiddis@amazon.com>
commit 9c97cfb
Author: Tomoyuki MORITA <moritato@amazon.com>
Date: Tue Oct 14 08:36:43 2025 -0700
Add JSON_EXTRACT_ALL internal function for Calcite PPL (opensearch-project#4489)
* Add JSON_EXTRACT_ALL internal function for Calcite PPL
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
* Address comments
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
* Minor fix
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
---------
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
commit 89dbc31
Author: Lantao Jin <ltjin@amazon.com>
Date: Tue Oct 14 18:24:52 2025 +0800
Check server status before starting Prometheus (opensearch-project#4537)
* Check server status before starting Prometheus
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* Change to func call
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* Fix doc
Signed-off-by: Lantao Jin <ltjin@amazon.com>
---------
Signed-off-by: Lantao Jin <ltjin@amazon.com>
commit fe62472
Author: Lantao Jin <ltjin@amazon.com>
Date: Tue Oct 14 18:10:27 2025 +0800
Update request builder after pushdown sort into agg buckets (opensearch-project#4541)
Signed-off-by: Lantao Jin <ltjin@amazon.com>
commit 42a415f
Author: qianheng <qianheng@amazon.com>
Date: Tue Oct 14 17:42:45 2025 +0800
Including metadata fields type when doing agg/filter script push down (opensearch-project#4522)
* Including metadata fields type when doing agg/filter script push down
Signed-off-by: Heng Qian <qianheng@amazon.com>
* Fix IT
Signed-off-by: Heng Qian <qianheng@amazon.com>
---------
Signed-off-by: Heng Qian <qianheng@amazon.com>
commit 8de0386
Author: Xinyuan Lu <xinyual@amazon.com>
Date: Tue Oct 14 16:41:08 2025 +0800
Fix percentile bug (opensearch-project#4539)
* fix percentile bug
Signed-off-by: xinyual <xinyual@amazon.com>
* add IT
Signed-off-by: xinyual <xinyual@amazon.com>
* optimize it
Signed-off-by: xinyual <xinyual@amazon.com>
---------
Signed-off-by: xinyual <xinyual@amazon.com>
commit de2fdc8
Author: Lantao Jin <ltjin@amazon.com>
Date: Tue Oct 14 12:29:03 2025 +0800
[FollowUp] Set 0 and negative value of subsearch.maxout as unlimited (opensearch-project#4534)
* [FollowUp] Set 0 and negative value of subsearch.maxout as unlimited
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* fix doctest
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* Fix conflicts
Signed-off-by: Lantao Jin <ltjin@amazon.com>
---------
Signed-off-by: Lantao Jin <ltjin@amazon.com>
commit 977b7ab
Author: Simeon Widdis <sawiddis@gmail.com>
Date: Mon Oct 13 20:23:10 2025 -0700
Update stalled action (opensearch-project#4485)
commit fddbb70
Author: Lantao Jin <ltjin@amazon.com>
Date: Tue Oct 14 10:23:12 2025 +0800
Add configurable sytem limitations for `subsearch` and `join` command (opensearch-project#4501)
* Add configurable sytem limitations for subsearch and join command
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* Fix IT
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* typo
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* fix IT
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* remove rollback in doc
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* address comments
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* fix typo
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* Fix IT
Signed-off-by: Lantao Jin <ltjin@amazon.com>
---------
Signed-off-by: Lantao Jin <ltjin@amazon.com>
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancementNew feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@LantaoJin@anasalkouz@penghuo@Swiddis
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Use _doc + _shard_doc as sort tiebreaker to get better performance - #4569

Merged
penghuo merged 2 commits into
opensearch-project:mainfrom
LantaoJin:pr/sort_tiebreaker
Oct 15, 2025
Merged

Use _doc + _shard_doc as sort tiebreaker to get better performance#4569
penghuo merged 2 commits into
opensearch-project:mainfrom
LantaoJin:pr/sort_tiebreaker

Conversation

@LantaoJin

@LantaoJinLantaoJin commented Oct 15, 2025

Copy link
Copy Markdown
Member

Description

Before #4378, the sort in PIT search is
case 1: if no sort field specified, sort by _doc + _id (+ means "then"). (❎ could cause high memory issue)
case 2: if sort fields specified, sort by fields. (❎ paged results could miss or duplicate hits)
case 3: if sort fields specified and query contains a filter, sort by _doc. (❎ paged results could miss or duplicate hits)

#4378 added the _shard_doc as sort tiebreaker with
case 1: if no sort field specified, sort by _shard_doc. (❎ performance regression)
case 2: if sort fields specified, sort by fields + _shard_doc.(❎ lower performance on low cardinality field)

#4435 found performance regression in case 1 and partially revert the changes to
case 1: if no sort field specified, sort by _doc + _id. (❎ could cause high memory issue)
case 2: if sort fields specified, sort by fields. (❎ paged results could miss or duplicate hits)

After this PR, we change the sort in PIT search to
case 1: if no sort field specified, sort by _doc + _shard_doc. ✅
case 2: if sort fields specified, sort by fields + _doc + _shard_doc.✅

RCA of performance regression:
_shard_doc is not a stored field in index which will be generated in runtime when comparison. Computing _shard_doc per document is a high cost operation. But sorting by _doc then _shard_doc only generates _shard_doc when the _doc values are conflicted.
Even in the case of user specified sort fields, we should sort by fields then _doc then _shard_doc to reduce the computing of _shard_doc. For example, if the sort field is a low cardinality field, e.g. gender, sorting by gender then _doc then _shard_doc generates _shard_doc for comparison only if values of gender and _doc are both conflicted.

This PR is no needed to backport to 2.19-dev since shard_doc feature is only available since OS 3.3.0

Related Issues

Resolves #[Issue number to be closed when this PR is merged]

Check List

  • New functionality includes testing.
  • New functionality has been documented.
  • New functionality has javadoc added.
  • New functionality has a user manual doc added.
  • New PPL command checklist all confirmed.
  • API changes companion pull request created.
  • Commits are signed per the DCO using --signoff or -s.
  • Public documentation issue/PR created.

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.
For more information on following Developer Certificate of Origin and signing off your commits, please check here.

Signed-off-by: Lantao Jin <ltjin@amazon.com>
@LantaoJinLantaoJin added the enhancement New feature or request label Oct 15, 2025
@LantaoJinLantaoJin changed the title Use _shard_doc as sort tiebreaker to get better performanceUse _doc + _shard_doc as sort tiebreaker to get better performanceOct 15, 2025
Signed-off-by: Lantao Jin <ltjin@amazon.com>
// Workaround to preserve sort location more exactly,
// see https://github.com/opensearch-project/sql/pull/3061
this.sourceBuilder.sort(METADATA_FIELD_ID, ASC);
this.sourceBuilder.sort(SortBuilders.shardDocSort());

@SwiddisSwiddisOct 15, 2025

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Does it matter if we duplicate fields in the sorting list? We could simplify/remove the below else logic by just always appending this, I would expect Lucene to optimize it in the background but I haven't measured it.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

not sure will Lucene optimize duplicated fields or _doc in sorting, but for sure the duplicated _shard_doc is not allowed in OpenSearch Core. It is no harmful for restricted checker here.

@anasalkouz

Copy link
Copy Markdown
Member

Can you share the performance benchmark for the 3 approaches?

@LantaoJin

LantaoJin commented Oct 15, 2025

Copy link
Copy Markdown
MemberAuthor

Can you share the performance benchmark for the 3 approaches?

I haven't run the benchmark, the RCA was made by reading the code of Luence and OS shard_doc feature.

The performance of _doc then _shard_doc is same as _doc then _id, provided by @ahkcs on Oct 1st. (the case 1)

For case 2, the current fields + _doc + _shard_doc is an further optimization upon fields + _shard_doc based on above benchmark result with inference.

Will rerun some benchmark to double confirm.

@penghuo
penghuo merged commit 3388dc7 into opensearch-project:mainOct 15, 2025
38 checks passed
ykmr1224 added a commit to ykmr1224/sql that referenced this pull request Oct 15, 2025
commit cba8d02
Author: Tomoyuki MORITA <moritato@amazon.com>
Date: Wed Oct 15 13:08:05 2025 -0700
Add MAP_APPEND internal function to Calcite PPL (opensearch-project#4515)
* Add MAP_APPEND internal function to Calcite PPL
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
* Minor fix
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
* Address comment
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
* Rebase and fix IT issue
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
---------
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
commit 3388dc7
Author: Lantao Jin <ltjin@amazon.com>
Date: Thu Oct 16 01:45:29 2025 +0800
Use `_doc` + `_shard_doc` as sort tiebreaker to get better performance (opensearch-project#4569)
* Use _shard_doc as sort tiebreaker
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* _doc as a part of tie-breaker have better performance
Signed-off-by: Lantao Jin <ltjin@amazon.com>
---------
Signed-off-by: Lantao Jin <ltjin@amazon.com>
commit 5630119
Author: qianheng <qianheng@amazon.com>
Date: Wed Oct 15 16:40:41 2025 +0800
Fix sort push down into agg after project already pushed (opensearch-project#4546)
* Fix sort push down into agg
Signed-off-by: Heng Qian <qianheng@amazon.com>
* Change some json files to yaml format
Signed-off-by: Heng Qian <qianheng@amazon.com>
---------
Signed-off-by: Heng Qian <qianheng@amazon.com>
commit 1e62fba
Author: Tomoyuki MORITA <moritato@amazon.com>
Date: Tue Oct 14 17:20:38 2025 -0700
Fix JsonExtractAllFunctionIT failure (opensearch-project#4556)
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
commit 02ee33e
Author: Kai Huang <105710027+ahkcs@users.noreply.github.com>
Date: Tue Oct 14 14:28:53 2025 -0700
Add more examples to the `where` command doc (opensearch-project#4457)
Co-authored-by: Manasvini B S <manasvis@amazon.com>
commit 0b7e86c
Author: Jialiang Liang <jiallian@amazon.com>
Date: Tue Oct 14 10:46:01 2025 -0700
[Enhancement] Error handling for illegal character usage in java regex named capture group (opensearch-project#4434)
Co-authored-by: Simeon Widdis <sawiddis@amazon.com>
commit 9c97cfb
Author: Tomoyuki MORITA <moritato@amazon.com>
Date: Tue Oct 14 08:36:43 2025 -0700
Add JSON_EXTRACT_ALL internal function for Calcite PPL (opensearch-project#4489)
* Add JSON_EXTRACT_ALL internal function for Calcite PPL
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
* Address comments
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
* Minor fix
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
---------
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
commit 89dbc31
Author: Lantao Jin <ltjin@amazon.com>
Date: Tue Oct 14 18:24:52 2025 +0800
Check server status before starting Prometheus (opensearch-project#4537)
* Check server status before starting Prometheus
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* Change to func call
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* Fix doc
Signed-off-by: Lantao Jin <ltjin@amazon.com>
---------
Signed-off-by: Lantao Jin <ltjin@amazon.com>
commit fe62472
Author: Lantao Jin <ltjin@amazon.com>
Date: Tue Oct 14 18:10:27 2025 +0800
Update request builder after pushdown sort into agg buckets (opensearch-project#4541)
Signed-off-by: Lantao Jin <ltjin@amazon.com>
commit 42a415f
Author: qianheng <qianheng@amazon.com>
Date: Tue Oct 14 17:42:45 2025 +0800
Including metadata fields type when doing agg/filter script push down (opensearch-project#4522)
* Including metadata fields type when doing agg/filter script push down
Signed-off-by: Heng Qian <qianheng@amazon.com>
* Fix IT
Signed-off-by: Heng Qian <qianheng@amazon.com>
---------
Signed-off-by: Heng Qian <qianheng@amazon.com>
commit 8de0386
Author: Xinyuan Lu <xinyual@amazon.com>
Date: Tue Oct 14 16:41:08 2025 +0800
Fix percentile bug (opensearch-project#4539)
* fix percentile bug
Signed-off-by: xinyual <xinyual@amazon.com>
* add IT
Signed-off-by: xinyual <xinyual@amazon.com>
* optimize it
Signed-off-by: xinyual <xinyual@amazon.com>
---------
Signed-off-by: xinyual <xinyual@amazon.com>
commit de2fdc8
Author: Lantao Jin <ltjin@amazon.com>
Date: Tue Oct 14 12:29:03 2025 +0800
[FollowUp] Set 0 and negative value of subsearch.maxout as unlimited (opensearch-project#4534)
* [FollowUp] Set 0 and negative value of subsearch.maxout as unlimited
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* fix doctest
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* Fix conflicts
Signed-off-by: Lantao Jin <ltjin@amazon.com>
---------
Signed-off-by: Lantao Jin <ltjin@amazon.com>
commit 977b7ab
Author: Simeon Widdis <sawiddis@gmail.com>
Date: Mon Oct 13 20:23:10 2025 -0700
Update stalled action (opensearch-project#4485)
commit fddbb70
Author: Lantao Jin <ltjin@amazon.com>
Date: Tue Oct 14 10:23:12 2025 +0800
Add configurable sytem limitations for `subsearch` and `join` command (opensearch-project#4501)
* Add configurable sytem limitations for subsearch and join command
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* Fix IT
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* typo
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* fix IT
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* remove rollback in doc
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* address comments
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* fix typo
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* Fix IT
Signed-off-by: Lantao Jin <ltjin@amazon.com>
---------
Signed-off-by: Lantao Jin <ltjin@amazon.com>
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancementNew feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@LantaoJin@anasalkouz@penghuo@Swiddis
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Use _doc + _shard_doc as sort tiebreaker to get better performance - #4569

Merged
penghuo merged 2 commits into
opensearch-project:mainfrom
LantaoJin:pr/sort_tiebreaker
Oct 15, 2025
Merged

Use _doc + _shard_doc as sort tiebreaker to get better performance#4569
penghuo merged 2 commits into
opensearch-project:mainfrom
LantaoJin:pr/sort_tiebreaker

Conversation

@LantaoJin

@LantaoJinLantaoJin commented Oct 15, 2025

Copy link
Copy Markdown
Member

Description

Before #4378, the sort in PIT search is
case 1: if no sort field specified, sort by _doc + _id (+ means "then"). (❎ could cause high memory issue)
case 2: if sort fields specified, sort by fields. (❎ paged results could miss or duplicate hits)
case 3: if sort fields specified and query contains a filter, sort by _doc. (❎ paged results could miss or duplicate hits)

#4378 added the _shard_doc as sort tiebreaker with
case 1: if no sort field specified, sort by _shard_doc. (❎ performance regression)
case 2: if sort fields specified, sort by fields + _shard_doc.(❎ lower performance on low cardinality field)

#4435 found performance regression in case 1 and partially revert the changes to
case 1: if no sort field specified, sort by _doc + _id. (❎ could cause high memory issue)
case 2: if sort fields specified, sort by fields. (❎ paged results could miss or duplicate hits)

After this PR, we change the sort in PIT search to
case 1: if no sort field specified, sort by _doc + _shard_doc. ✅
case 2: if sort fields specified, sort by fields + _doc + _shard_doc.✅

RCA of performance regression:
_shard_doc is not a stored field in index which will be generated in runtime when comparison. Computing _shard_doc per document is a high cost operation. But sorting by _doc then _shard_doc only generates _shard_doc when the _doc values are conflicted.
Even in the case of user specified sort fields, we should sort by fields then _doc then _shard_doc to reduce the computing of _shard_doc. For example, if the sort field is a low cardinality field, e.g. gender, sorting by gender then _doc then _shard_doc generates _shard_doc for comparison only if values of gender and _doc are both conflicted.

This PR is no needed to backport to 2.19-dev since shard_doc feature is only available since OS 3.3.0

Related Issues

Resolves #[Issue number to be closed when this PR is merged]

Check List

  • New functionality includes testing.
  • New functionality has been documented.
  • New functionality has javadoc added.
  • New functionality has a user manual doc added.
  • New PPL command checklist all confirmed.
  • API changes companion pull request created.
  • Commits are signed per the DCO using --signoff or -s.
  • Public documentation issue/PR created.

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.
For more information on following Developer Certificate of Origin and signing off your commits, please check here.

Signed-off-by: Lantao Jin <ltjin@amazon.com>
@LantaoJinLantaoJin added the enhancement New feature or request label Oct 15, 2025
@LantaoJinLantaoJin changed the title Use _shard_doc as sort tiebreaker to get better performanceUse _doc + _shard_doc as sort tiebreaker to get better performanceOct 15, 2025
Signed-off-by: Lantao Jin <ltjin@amazon.com>
// Workaround to preserve sort location more exactly,
// see https://github.com/opensearch-project/sql/pull/3061
this.sourceBuilder.sort(METADATA_FIELD_ID, ASC);
this.sourceBuilder.sort(SortBuilders.shardDocSort());

@SwiddisSwiddisOct 15, 2025

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Does it matter if we duplicate fields in the sorting list? We could simplify/remove the below else logic by just always appending this, I would expect Lucene to optimize it in the background but I haven't measured it.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

not sure will Lucene optimize duplicated fields or _doc in sorting, but for sure the duplicated _shard_doc is not allowed in OpenSearch Core. It is no harmful for restricted checker here.

@anasalkouz

Copy link
Copy Markdown
Member

Can you share the performance benchmark for the 3 approaches?

@LantaoJin

LantaoJin commented Oct 15, 2025

Copy link
Copy Markdown
MemberAuthor

Can you share the performance benchmark for the 3 approaches?

I haven't run the benchmark, the RCA was made by reading the code of Luence and OS shard_doc feature.

The performance of _doc then _shard_doc is same as _doc then _id, provided by @ahkcs on Oct 1st. (the case 1)

For case 2, the current fields + _doc + _shard_doc is an further optimization upon fields + _shard_doc based on above benchmark result with inference.

Will rerun some benchmark to double confirm.

@penghuo
penghuo merged commit 3388dc7 into opensearch-project:mainOct 15, 2025
38 checks passed
ykmr1224 added a commit to ykmr1224/sql that referenced this pull request Oct 15, 2025
commit cba8d02
Author: Tomoyuki MORITA <moritato@amazon.com>
Date: Wed Oct 15 13:08:05 2025 -0700
Add MAP_APPEND internal function to Calcite PPL (opensearch-project#4515)
* Add MAP_APPEND internal function to Calcite PPL
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
* Minor fix
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
* Address comment
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
* Rebase and fix IT issue
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
---------
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
commit 3388dc7
Author: Lantao Jin <ltjin@amazon.com>
Date: Thu Oct 16 01:45:29 2025 +0800
Use `_doc` + `_shard_doc` as sort tiebreaker to get better performance (opensearch-project#4569)
* Use _shard_doc as sort tiebreaker
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* _doc as a part of tie-breaker have better performance
Signed-off-by: Lantao Jin <ltjin@amazon.com>
---------
Signed-off-by: Lantao Jin <ltjin@amazon.com>
commit 5630119
Author: qianheng <qianheng@amazon.com>
Date: Wed Oct 15 16:40:41 2025 +0800
Fix sort push down into agg after project already pushed (opensearch-project#4546)
* Fix sort push down into agg
Signed-off-by: Heng Qian <qianheng@amazon.com>
* Change some json files to yaml format
Signed-off-by: Heng Qian <qianheng@amazon.com>
---------
Signed-off-by: Heng Qian <qianheng@amazon.com>
commit 1e62fba
Author: Tomoyuki MORITA <moritato@amazon.com>
Date: Tue Oct 14 17:20:38 2025 -0700
Fix JsonExtractAllFunctionIT failure (opensearch-project#4556)
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
commit 02ee33e
Author: Kai Huang <105710027+ahkcs@users.noreply.github.com>
Date: Tue Oct 14 14:28:53 2025 -0700
Add more examples to the `where` command doc (opensearch-project#4457)
Co-authored-by: Manasvini B S <manasvis@amazon.com>
commit 0b7e86c
Author: Jialiang Liang <jiallian@amazon.com>
Date: Tue Oct 14 10:46:01 2025 -0700
[Enhancement] Error handling for illegal character usage in java regex named capture group (opensearch-project#4434)
Co-authored-by: Simeon Widdis <sawiddis@amazon.com>
commit 9c97cfb
Author: Tomoyuki MORITA <moritato@amazon.com>
Date: Tue Oct 14 08:36:43 2025 -0700
Add JSON_EXTRACT_ALL internal function for Calcite PPL (opensearch-project#4489)
* Add JSON_EXTRACT_ALL internal function for Calcite PPL
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
* Address comments
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
* Minor fix
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
---------
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
commit 89dbc31
Author: Lantao Jin <ltjin@amazon.com>
Date: Tue Oct 14 18:24:52 2025 +0800
Check server status before starting Prometheus (opensearch-project#4537)
* Check server status before starting Prometheus
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* Change to func call
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* Fix doc
Signed-off-by: Lantao Jin <ltjin@amazon.com>
---------
Signed-off-by: Lantao Jin <ltjin@amazon.com>
commit fe62472
Author: Lantao Jin <ltjin@amazon.com>
Date: Tue Oct 14 18:10:27 2025 +0800
Update request builder after pushdown sort into agg buckets (opensearch-project#4541)
Signed-off-by: Lantao Jin <ltjin@amazon.com>
commit 42a415f
Author: qianheng <qianheng@amazon.com>
Date: Tue Oct 14 17:42:45 2025 +0800
Including metadata fields type when doing agg/filter script push down (opensearch-project#4522)
* Including metadata fields type when doing agg/filter script push down
Signed-off-by: Heng Qian <qianheng@amazon.com>
* Fix IT
Signed-off-by: Heng Qian <qianheng@amazon.com>
---------
Signed-off-by: Heng Qian <qianheng@amazon.com>
commit 8de0386
Author: Xinyuan Lu <xinyual@amazon.com>
Date: Tue Oct 14 16:41:08 2025 +0800
Fix percentile bug (opensearch-project#4539)
* fix percentile bug
Signed-off-by: xinyual <xinyual@amazon.com>
* add IT
Signed-off-by: xinyual <xinyual@amazon.com>
* optimize it
Signed-off-by: xinyual <xinyual@amazon.com>
---------
Signed-off-by: xinyual <xinyual@amazon.com>
commit de2fdc8
Author: Lantao Jin <ltjin@amazon.com>
Date: Tue Oct 14 12:29:03 2025 +0800
[FollowUp] Set 0 and negative value of subsearch.maxout as unlimited (opensearch-project#4534)
* [FollowUp] Set 0 and negative value of subsearch.maxout as unlimited
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* fix doctest
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* Fix conflicts
Signed-off-by: Lantao Jin <ltjin@amazon.com>
---------
Signed-off-by: Lantao Jin <ltjin@amazon.com>
commit 977b7ab
Author: Simeon Widdis <sawiddis@gmail.com>
Date: Mon Oct 13 20:23:10 2025 -0700
Update stalled action (opensearch-project#4485)
commit fddbb70
Author: Lantao Jin <ltjin@amazon.com>
Date: Tue Oct 14 10:23:12 2025 +0800
Add configurable sytem limitations for `subsearch` and `join` command (opensearch-project#4501)
* Add configurable sytem limitations for subsearch and join command
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* Fix IT
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* typo
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* fix IT
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* remove rollback in doc
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* address comments
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* fix typo
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* Fix IT
Signed-off-by: Lantao Jin <ltjin@amazon.com>
---------
Signed-off-by: Lantao Jin <ltjin@amazon.com>
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancementNew feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@LantaoJin@anasalkouz@penghuo@Swiddis
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Use _doc + _shard_doc as sort tiebreaker to get better performance - #4569

Merged
penghuo merged 2 commits into
opensearch-project:mainfrom
LantaoJin:pr/sort_tiebreaker
Oct 15, 2025
Merged

Use _doc + _shard_doc as sort tiebreaker to get better performance#4569
penghuo merged 2 commits into
opensearch-project:mainfrom
LantaoJin:pr/sort_tiebreaker

Conversation

@LantaoJin

@LantaoJinLantaoJin commented Oct 15, 2025

Copy link
Copy Markdown
Member

Description

Before #4378, the sort in PIT search is
case 1: if no sort field specified, sort by _doc + _id (+ means "then"). (❎ could cause high memory issue)
case 2: if sort fields specified, sort by fields. (❎ paged results could miss or duplicate hits)
case 3: if sort fields specified and query contains a filter, sort by _doc. (❎ paged results could miss or duplicate hits)

#4378 added the _shard_doc as sort tiebreaker with
case 1: if no sort field specified, sort by _shard_doc. (❎ performance regression)
case 2: if sort fields specified, sort by fields + _shard_doc.(❎ lower performance on low cardinality field)

#4435 found performance regression in case 1 and partially revert the changes to
case 1: if no sort field specified, sort by _doc + _id. (❎ could cause high memory issue)
case 2: if sort fields specified, sort by fields. (❎ paged results could miss or duplicate hits)

After this PR, we change the sort in PIT search to
case 1: if no sort field specified, sort by _doc + _shard_doc. ✅
case 2: if sort fields specified, sort by fields + _doc + _shard_doc.✅

RCA of performance regression:
_shard_doc is not a stored field in index which will be generated in runtime when comparison. Computing _shard_doc per document is a high cost operation. But sorting by _doc then _shard_doc only generates _shard_doc when the _doc values are conflicted.
Even in the case of user specified sort fields, we should sort by fields then _doc then _shard_doc to reduce the computing of _shard_doc. For example, if the sort field is a low cardinality field, e.g. gender, sorting by gender then _doc then _shard_doc generates _shard_doc for comparison only if values of gender and _doc are both conflicted.

This PR is no needed to backport to 2.19-dev since shard_doc feature is only available since OS 3.3.0

Related Issues

Resolves #[Issue number to be closed when this PR is merged]

Check List

  • New functionality includes testing.
  • New functionality has been documented.
  • New functionality has javadoc added.
  • New functionality has a user manual doc added.
  • New PPL command checklist all confirmed.
  • API changes companion pull request created.
  • Commits are signed per the DCO using --signoff or -s.
  • Public documentation issue/PR created.

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.
For more information on following Developer Certificate of Origin and signing off your commits, please check here.

Signed-off-by: Lantao Jin <ltjin@amazon.com>
@LantaoJinLantaoJin added the enhancement New feature or request label Oct 15, 2025
@LantaoJinLantaoJin changed the title Use _shard_doc as sort tiebreaker to get better performanceUse _doc + _shard_doc as sort tiebreaker to get better performanceOct 15, 2025
Signed-off-by: Lantao Jin <ltjin@amazon.com>
// Workaround to preserve sort location more exactly,
// see https://github.com/opensearch-project/sql/pull/3061
this.sourceBuilder.sort(METADATA_FIELD_ID, ASC);
this.sourceBuilder.sort(SortBuilders.shardDocSort());

@SwiddisSwiddisOct 15, 2025

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Does it matter if we duplicate fields in the sorting list? We could simplify/remove the below else logic by just always appending this, I would expect Lucene to optimize it in the background but I haven't measured it.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

not sure will Lucene optimize duplicated fields or _doc in sorting, but for sure the duplicated _shard_doc is not allowed in OpenSearch Core. It is no harmful for restricted checker here.

@anasalkouz

Copy link
Copy Markdown
Member

Can you share the performance benchmark for the 3 approaches?

@LantaoJin

LantaoJin commented Oct 15, 2025

Copy link
Copy Markdown
MemberAuthor

Can you share the performance benchmark for the 3 approaches?

I haven't run the benchmark, the RCA was made by reading the code of Luence and OS shard_doc feature.

The performance of _doc then _shard_doc is same as _doc then _id, provided by @ahkcs on Oct 1st. (the case 1)

For case 2, the current fields + _doc + _shard_doc is an further optimization upon fields + _shard_doc based on above benchmark result with inference.

Will rerun some benchmark to double confirm.

@penghuo
penghuo merged commit 3388dc7 into opensearch-project:mainOct 15, 2025
38 checks passed
ykmr1224 added a commit to ykmr1224/sql that referenced this pull request Oct 15, 2025
commit cba8d02
Author: Tomoyuki MORITA <moritato@amazon.com>
Date: Wed Oct 15 13:08:05 2025 -0700
Add MAP_APPEND internal function to Calcite PPL (opensearch-project#4515)
* Add MAP_APPEND internal function to Calcite PPL
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
* Minor fix
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
* Address comment
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
* Rebase and fix IT issue
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
---------
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
commit 3388dc7
Author: Lantao Jin <ltjin@amazon.com>
Date: Thu Oct 16 01:45:29 2025 +0800
Use `_doc` + `_shard_doc` as sort tiebreaker to get better performance (opensearch-project#4569)
* Use _shard_doc as sort tiebreaker
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* _doc as a part of tie-breaker have better performance
Signed-off-by: Lantao Jin <ltjin@amazon.com>
---------
Signed-off-by: Lantao Jin <ltjin@amazon.com>
commit 5630119
Author: qianheng <qianheng@amazon.com>
Date: Wed Oct 15 16:40:41 2025 +0800
Fix sort push down into agg after project already pushed (opensearch-project#4546)
* Fix sort push down into agg
Signed-off-by: Heng Qian <qianheng@amazon.com>
* Change some json files to yaml format
Signed-off-by: Heng Qian <qianheng@amazon.com>
---------
Signed-off-by: Heng Qian <qianheng@amazon.com>
commit 1e62fba
Author: Tomoyuki MORITA <moritato@amazon.com>
Date: Tue Oct 14 17:20:38 2025 -0700
Fix JsonExtractAllFunctionIT failure (opensearch-project#4556)
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
commit 02ee33e
Author: Kai Huang <105710027+ahkcs@users.noreply.github.com>
Date: Tue Oct 14 14:28:53 2025 -0700
Add more examples to the `where` command doc (opensearch-project#4457)
Co-authored-by: Manasvini B S <manasvis@amazon.com>
commit 0b7e86c
Author: Jialiang Liang <jiallian@amazon.com>
Date: Tue Oct 14 10:46:01 2025 -0700
[Enhancement] Error handling for illegal character usage in java regex named capture group (opensearch-project#4434)
Co-authored-by: Simeon Widdis <sawiddis@amazon.com>
commit 9c97cfb
Author: Tomoyuki MORITA <moritato@amazon.com>
Date: Tue Oct 14 08:36:43 2025 -0700
Add JSON_EXTRACT_ALL internal function for Calcite PPL (opensearch-project#4489)
* Add JSON_EXTRACT_ALL internal function for Calcite PPL
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
* Address comments
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
* Minor fix
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
---------
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
commit 89dbc31
Author: Lantao Jin <ltjin@amazon.com>
Date: Tue Oct 14 18:24:52 2025 +0800
Check server status before starting Prometheus (opensearch-project#4537)
* Check server status before starting Prometheus
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* Change to func call
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* Fix doc
Signed-off-by: Lantao Jin <ltjin@amazon.com>
---------
Signed-off-by: Lantao Jin <ltjin@amazon.com>
commit fe62472
Author: Lantao Jin <ltjin@amazon.com>
Date: Tue Oct 14 18:10:27 2025 +0800
Update request builder after pushdown sort into agg buckets (opensearch-project#4541)
Signed-off-by: Lantao Jin <ltjin@amazon.com>
commit 42a415f
Author: qianheng <qianheng@amazon.com>
Date: Tue Oct 14 17:42:45 2025 +0800
Including metadata fields type when doing agg/filter script push down (opensearch-project#4522)
* Including metadata fields type when doing agg/filter script push down
Signed-off-by: Heng Qian <qianheng@amazon.com>
* Fix IT
Signed-off-by: Heng Qian <qianheng@amazon.com>
---------
Signed-off-by: Heng Qian <qianheng@amazon.com>
commit 8de0386
Author: Xinyuan Lu <xinyual@amazon.com>
Date: Tue Oct 14 16:41:08 2025 +0800
Fix percentile bug (opensearch-project#4539)
* fix percentile bug
Signed-off-by: xinyual <xinyual@amazon.com>
* add IT
Signed-off-by: xinyual <xinyual@amazon.com>
* optimize it
Signed-off-by: xinyual <xinyual@amazon.com>
---------
Signed-off-by: xinyual <xinyual@amazon.com>
commit de2fdc8
Author: Lantao Jin <ltjin@amazon.com>
Date: Tue Oct 14 12:29:03 2025 +0800
[FollowUp] Set 0 and negative value of subsearch.maxout as unlimited (opensearch-project#4534)
* [FollowUp] Set 0 and negative value of subsearch.maxout as unlimited
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* fix doctest
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* Fix conflicts
Signed-off-by: Lantao Jin <ltjin@amazon.com>
---------
Signed-off-by: Lantao Jin <ltjin@amazon.com>
commit 977b7ab
Author: Simeon Widdis <sawiddis@gmail.com>
Date: Mon Oct 13 20:23:10 2025 -0700
Update stalled action (opensearch-project#4485)
commit fddbb70
Author: Lantao Jin <ltjin@amazon.com>
Date: Tue Oct 14 10:23:12 2025 +0800
Add configurable sytem limitations for `subsearch` and `join` command (opensearch-project#4501)
* Add configurable sytem limitations for subsearch and join command
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* Fix IT
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* typo
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* fix IT
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* remove rollback in doc
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* address comments
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* fix typo
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* Fix IT
Signed-off-by: Lantao Jin <ltjin@amazon.com>
---------
Signed-off-by: Lantao Jin <ltjin@amazon.com>
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancementNew feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@LantaoJin@anasalkouz@penghuo@Swiddis
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Use _doc + _shard_doc as sort tiebreaker to get better performance - #4569

Merged
penghuo merged 2 commits into
opensearch-project:mainfrom
LantaoJin:pr/sort_tiebreaker
Oct 15, 2025
Merged

Use _doc + _shard_doc as sort tiebreaker to get better performance#4569
penghuo merged 2 commits into
opensearch-project:mainfrom
LantaoJin:pr/sort_tiebreaker

Conversation

@LantaoJin

@LantaoJinLantaoJin commented Oct 15, 2025

Copy link
Copy Markdown
Member

Description

Before #4378, the sort in PIT search is
case 1: if no sort field specified, sort by _doc + _id (+ means "then"). (❎ could cause high memory issue)
case 2: if sort fields specified, sort by fields. (❎ paged results could miss or duplicate hits)
case 3: if sort fields specified and query contains a filter, sort by _doc. (❎ paged results could miss or duplicate hits)

#4378 added the _shard_doc as sort tiebreaker with
case 1: if no sort field specified, sort by _shard_doc. (❎ performance regression)
case 2: if sort fields specified, sort by fields + _shard_doc.(❎ lower performance on low cardinality field)

#4435 found performance regression in case 1 and partially revert the changes to
case 1: if no sort field specified, sort by _doc + _id. (❎ could cause high memory issue)
case 2: if sort fields specified, sort by fields. (❎ paged results could miss or duplicate hits)

After this PR, we change the sort in PIT search to
case 1: if no sort field specified, sort by _doc + _shard_doc. ✅
case 2: if sort fields specified, sort by fields + _doc + _shard_doc.✅

RCA of performance regression:
_shard_doc is not a stored field in index which will be generated in runtime when comparison. Computing _shard_doc per document is a high cost operation. But sorting by _doc then _shard_doc only generates _shard_doc when the _doc values are conflicted.
Even in the case of user specified sort fields, we should sort by fields then _doc then _shard_doc to reduce the computing of _shard_doc. For example, if the sort field is a low cardinality field, e.g. gender, sorting by gender then _doc then _shard_doc generates _shard_doc for comparison only if values of gender and _doc are both conflicted.

This PR is no needed to backport to 2.19-dev since shard_doc feature is only available since OS 3.3.0

Related Issues

Resolves #[Issue number to be closed when this PR is merged]

Check List

  • New functionality includes testing.
  • New functionality has been documented.
  • New functionality has javadoc added.
  • New functionality has a user manual doc added.
  • New PPL command checklist all confirmed.
  • API changes companion pull request created.
  • Commits are signed per the DCO using --signoff or -s.
  • Public documentation issue/PR created.

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.
For more information on following Developer Certificate of Origin and signing off your commits, please check here.

Signed-off-by: Lantao Jin <ltjin@amazon.com>
@LantaoJinLantaoJin added the enhancement New feature or request label Oct 15, 2025
@LantaoJinLantaoJin changed the title Use _shard_doc as sort tiebreaker to get better performanceUse _doc + _shard_doc as sort tiebreaker to get better performanceOct 15, 2025
Signed-off-by: Lantao Jin <ltjin@amazon.com>
// Workaround to preserve sort location more exactly,
// see https://github.com/opensearch-project/sql/pull/3061
this.sourceBuilder.sort(METADATA_FIELD_ID, ASC);
this.sourceBuilder.sort(SortBuilders.shardDocSort());

@SwiddisSwiddisOct 15, 2025

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Does it matter if we duplicate fields in the sorting list? We could simplify/remove the below else logic by just always appending this, I would expect Lucene to optimize it in the background but I haven't measured it.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

not sure will Lucene optimize duplicated fields or _doc in sorting, but for sure the duplicated _shard_doc is not allowed in OpenSearch Core. It is no harmful for restricted checker here.

@anasalkouz

Copy link
Copy Markdown
Member

Can you share the performance benchmark for the 3 approaches?

@LantaoJin

LantaoJin commented Oct 15, 2025

Copy link
Copy Markdown
MemberAuthor

Can you share the performance benchmark for the 3 approaches?

I haven't run the benchmark, the RCA was made by reading the code of Luence and OS shard_doc feature.

The performance of _doc then _shard_doc is same as _doc then _id, provided by @ahkcs on Oct 1st. (the case 1)

For case 2, the current fields + _doc + _shard_doc is an further optimization upon fields + _shard_doc based on above benchmark result with inference.

Will rerun some benchmark to double confirm.

@penghuo
penghuo merged commit 3388dc7 into opensearch-project:mainOct 15, 2025
38 checks passed
ykmr1224 added a commit to ykmr1224/sql that referenced this pull request Oct 15, 2025
commit cba8d02
Author: Tomoyuki MORITA <moritato@amazon.com>
Date: Wed Oct 15 13:08:05 2025 -0700
Add MAP_APPEND internal function to Calcite PPL (opensearch-project#4515)
* Add MAP_APPEND internal function to Calcite PPL
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
* Minor fix
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
* Address comment
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
* Rebase and fix IT issue
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
---------
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
commit 3388dc7
Author: Lantao Jin <ltjin@amazon.com>
Date: Thu Oct 16 01:45:29 2025 +0800
Use `_doc` + `_shard_doc` as sort tiebreaker to get better performance (opensearch-project#4569)
* Use _shard_doc as sort tiebreaker
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* _doc as a part of tie-breaker have better performance
Signed-off-by: Lantao Jin <ltjin@amazon.com>
---------
Signed-off-by: Lantao Jin <ltjin@amazon.com>
commit 5630119
Author: qianheng <qianheng@amazon.com>
Date: Wed Oct 15 16:40:41 2025 +0800
Fix sort push down into agg after project already pushed (opensearch-project#4546)
* Fix sort push down into agg
Signed-off-by: Heng Qian <qianheng@amazon.com>
* Change some json files to yaml format
Signed-off-by: Heng Qian <qianheng@amazon.com>
---------
Signed-off-by: Heng Qian <qianheng@amazon.com>
commit 1e62fba
Author: Tomoyuki MORITA <moritato@amazon.com>
Date: Tue Oct 14 17:20:38 2025 -0700
Fix JsonExtractAllFunctionIT failure (opensearch-project#4556)
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
commit 02ee33e
Author: Kai Huang <105710027+ahkcs@users.noreply.github.com>
Date: Tue Oct 14 14:28:53 2025 -0700
Add more examples to the `where` command doc (opensearch-project#4457)
Co-authored-by: Manasvini B S <manasvis@amazon.com>
commit 0b7e86c
Author: Jialiang Liang <jiallian@amazon.com>
Date: Tue Oct 14 10:46:01 2025 -0700
[Enhancement] Error handling for illegal character usage in java regex named capture group (opensearch-project#4434)
Co-authored-by: Simeon Widdis <sawiddis@amazon.com>
commit 9c97cfb
Author: Tomoyuki MORITA <moritato@amazon.com>
Date: Tue Oct 14 08:36:43 2025 -0700
Add JSON_EXTRACT_ALL internal function for Calcite PPL (opensearch-project#4489)
* Add JSON_EXTRACT_ALL internal function for Calcite PPL
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
* Address comments
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
* Minor fix
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
---------
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
commit 89dbc31
Author: Lantao Jin <ltjin@amazon.com>
Date: Tue Oct 14 18:24:52 2025 +0800
Check server status before starting Prometheus (opensearch-project#4537)
* Check server status before starting Prometheus
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* Change to func call
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* Fix doc
Signed-off-by: Lantao Jin <ltjin@amazon.com>
---------
Signed-off-by: Lantao Jin <ltjin@amazon.com>
commit fe62472
Author: Lantao Jin <ltjin@amazon.com>
Date: Tue Oct 14 18:10:27 2025 +0800
Update request builder after pushdown sort into agg buckets (opensearch-project#4541)
Signed-off-by: Lantao Jin <ltjin@amazon.com>
commit 42a415f
Author: qianheng <qianheng@amazon.com>
Date: Tue Oct 14 17:42:45 2025 +0800
Including metadata fields type when doing agg/filter script push down (opensearch-project#4522)
* Including metadata fields type when doing agg/filter script push down
Signed-off-by: Heng Qian <qianheng@amazon.com>
* Fix IT
Signed-off-by: Heng Qian <qianheng@amazon.com>
---------
Signed-off-by: Heng Qian <qianheng@amazon.com>
commit 8de0386
Author: Xinyuan Lu <xinyual@amazon.com>
Date: Tue Oct 14 16:41:08 2025 +0800
Fix percentile bug (opensearch-project#4539)
* fix percentile bug
Signed-off-by: xinyual <xinyual@amazon.com>
* add IT
Signed-off-by: xinyual <xinyual@amazon.com>
* optimize it
Signed-off-by: xinyual <xinyual@amazon.com>
---------
Signed-off-by: xinyual <xinyual@amazon.com>
commit de2fdc8
Author: Lantao Jin <ltjin@amazon.com>
Date: Tue Oct 14 12:29:03 2025 +0800
[FollowUp] Set 0 and negative value of subsearch.maxout as unlimited (opensearch-project#4534)
* [FollowUp] Set 0 and negative value of subsearch.maxout as unlimited
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* fix doctest
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* Fix conflicts
Signed-off-by: Lantao Jin <ltjin@amazon.com>
---------
Signed-off-by: Lantao Jin <ltjin@amazon.com>
commit 977b7ab
Author: Simeon Widdis <sawiddis@gmail.com>
Date: Mon Oct 13 20:23:10 2025 -0700
Update stalled action (opensearch-project#4485)
commit fddbb70
Author: Lantao Jin <ltjin@amazon.com>
Date: Tue Oct 14 10:23:12 2025 +0800
Add configurable sytem limitations for `subsearch` and `join` command (opensearch-project#4501)
* Add configurable sytem limitations for subsearch and join command
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* Fix IT
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* typo
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* fix IT
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* remove rollback in doc
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* address comments
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* fix typo
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* Fix IT
Signed-off-by: Lantao Jin <ltjin@amazon.com>
---------
Signed-off-by: Lantao Jin <ltjin@amazon.com>
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancementNew feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@LantaoJin@anasalkouz@penghuo@Swiddis
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Use _doc + _shard_doc as sort tiebreaker to get better performance - #4569

Merged
penghuo merged 2 commits into
opensearch-project:mainfrom
LantaoJin:pr/sort_tiebreaker
Oct 15, 2025
Merged

Use _doc + _shard_doc as sort tiebreaker to get better performance#4569
penghuo merged 2 commits into
opensearch-project:mainfrom
LantaoJin:pr/sort_tiebreaker

Conversation

@LantaoJin

@LantaoJinLantaoJin commented Oct 15, 2025

Copy link
Copy Markdown
Member

Description

Before #4378, the sort in PIT search is
case 1: if no sort field specified, sort by _doc + _id (+ means "then"). (❎ could cause high memory issue)
case 2: if sort fields specified, sort by fields. (❎ paged results could miss or duplicate hits)
case 3: if sort fields specified and query contains a filter, sort by _doc. (❎ paged results could miss or duplicate hits)

#4378 added the _shard_doc as sort tiebreaker with
case 1: if no sort field specified, sort by _shard_doc. (❎ performance regression)
case 2: if sort fields specified, sort by fields + _shard_doc.(❎ lower performance on low cardinality field)

#4435 found performance regression in case 1 and partially revert the changes to
case 1: if no sort field specified, sort by _doc + _id. (❎ could cause high memory issue)
case 2: if sort fields specified, sort by fields. (❎ paged results could miss or duplicate hits)

After this PR, we change the sort in PIT search to
case 1: if no sort field specified, sort by _doc + _shard_doc. ✅
case 2: if sort fields specified, sort by fields + _doc + _shard_doc.✅

RCA of performance regression:
_shard_doc is not a stored field in index which will be generated in runtime when comparison. Computing _shard_doc per document is a high cost operation. But sorting by _doc then _shard_doc only generates _shard_doc when the _doc values are conflicted.
Even in the case of user specified sort fields, we should sort by fields then _doc then _shard_doc to reduce the computing of _shard_doc. For example, if the sort field is a low cardinality field, e.g. gender, sorting by gender then _doc then _shard_doc generates _shard_doc for comparison only if values of gender and _doc are both conflicted.

This PR is no needed to backport to 2.19-dev since shard_doc feature is only available since OS 3.3.0

Related Issues

Resolves #[Issue number to be closed when this PR is merged]

Check List

  • New functionality includes testing.
  • New functionality has been documented.
  • New functionality has javadoc added.
  • New functionality has a user manual doc added.
  • New PPL command checklist all confirmed.
  • API changes companion pull request created.
  • Commits are signed per the DCO using --signoff or -s.
  • Public documentation issue/PR created.

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.
For more information on following Developer Certificate of Origin and signing off your commits, please check here.

Signed-off-by: Lantao Jin <ltjin@amazon.com>
@LantaoJinLantaoJin added the enhancement New feature or request label Oct 15, 2025
@LantaoJinLantaoJin changed the title Use _shard_doc as sort tiebreaker to get better performanceUse _doc + _shard_doc as sort tiebreaker to get better performanceOct 15, 2025
Signed-off-by: Lantao Jin <ltjin@amazon.com>
// Workaround to preserve sort location more exactly,
// see https://github.com/opensearch-project/sql/pull/3061
this.sourceBuilder.sort(METADATA_FIELD_ID, ASC);
this.sourceBuilder.sort(SortBuilders.shardDocSort());

@SwiddisSwiddisOct 15, 2025

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Does it matter if we duplicate fields in the sorting list? We could simplify/remove the below else logic by just always appending this, I would expect Lucene to optimize it in the background but I haven't measured it.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

not sure will Lucene optimize duplicated fields or _doc in sorting, but for sure the duplicated _shard_doc is not allowed in OpenSearch Core. It is no harmful for restricted checker here.

@anasalkouz

Copy link
Copy Markdown
Member

Can you share the performance benchmark for the 3 approaches?

@LantaoJin

LantaoJin commented Oct 15, 2025

Copy link
Copy Markdown
MemberAuthor

Can you share the performance benchmark for the 3 approaches?

I haven't run the benchmark, the RCA was made by reading the code of Luence and OS shard_doc feature.

The performance of _doc then _shard_doc is same as _doc then _id, provided by @ahkcs on Oct 1st. (the case 1)

For case 2, the current fields + _doc + _shard_doc is an further optimization upon fields + _shard_doc based on above benchmark result with inference.

Will rerun some benchmark to double confirm.

@penghuo
penghuo merged commit 3388dc7 into opensearch-project:mainOct 15, 2025
38 checks passed
ykmr1224 added a commit to ykmr1224/sql that referenced this pull request Oct 15, 2025
commit cba8d02
Author: Tomoyuki MORITA <moritato@amazon.com>
Date: Wed Oct 15 13:08:05 2025 -0700
Add MAP_APPEND internal function to Calcite PPL (opensearch-project#4515)
* Add MAP_APPEND internal function to Calcite PPL
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
* Minor fix
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
* Address comment
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
* Rebase and fix IT issue
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
---------
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
commit 3388dc7
Author: Lantao Jin <ltjin@amazon.com>
Date: Thu Oct 16 01:45:29 2025 +0800
Use `_doc` + `_shard_doc` as sort tiebreaker to get better performance (opensearch-project#4569)
* Use _shard_doc as sort tiebreaker
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* _doc as a part of tie-breaker have better performance
Signed-off-by: Lantao Jin <ltjin@amazon.com>
---------
Signed-off-by: Lantao Jin <ltjin@amazon.com>
commit 5630119
Author: qianheng <qianheng@amazon.com>
Date: Wed Oct 15 16:40:41 2025 +0800
Fix sort push down into agg after project already pushed (opensearch-project#4546)
* Fix sort push down into agg
Signed-off-by: Heng Qian <qianheng@amazon.com>
* Change some json files to yaml format
Signed-off-by: Heng Qian <qianheng@amazon.com>
---------
Signed-off-by: Heng Qian <qianheng@amazon.com>
commit 1e62fba
Author: Tomoyuki MORITA <moritato@amazon.com>
Date: Tue Oct 14 17:20:38 2025 -0700
Fix JsonExtractAllFunctionIT failure (opensearch-project#4556)
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
commit 02ee33e
Author: Kai Huang <105710027+ahkcs@users.noreply.github.com>
Date: Tue Oct 14 14:28:53 2025 -0700
Add more examples to the `where` command doc (opensearch-project#4457)
Co-authored-by: Manasvini B S <manasvis@amazon.com>
commit 0b7e86c
Author: Jialiang Liang <jiallian@amazon.com>
Date: Tue Oct 14 10:46:01 2025 -0700
[Enhancement] Error handling for illegal character usage in java regex named capture group (opensearch-project#4434)
Co-authored-by: Simeon Widdis <sawiddis@amazon.com>
commit 9c97cfb
Author: Tomoyuki MORITA <moritato@amazon.com>
Date: Tue Oct 14 08:36:43 2025 -0700
Add JSON_EXTRACT_ALL internal function for Calcite PPL (opensearch-project#4489)
* Add JSON_EXTRACT_ALL internal function for Calcite PPL
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
* Address comments
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
* Minor fix
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
---------
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
commit 89dbc31
Author: Lantao Jin <ltjin@amazon.com>
Date: Tue Oct 14 18:24:52 2025 +0800
Check server status before starting Prometheus (opensearch-project#4537)
* Check server status before starting Prometheus
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* Change to func call
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* Fix doc
Signed-off-by: Lantao Jin <ltjin@amazon.com>
---------
Signed-off-by: Lantao Jin <ltjin@amazon.com>
commit fe62472
Author: Lantao Jin <ltjin@amazon.com>
Date: Tue Oct 14 18:10:27 2025 +0800
Update request builder after pushdown sort into agg buckets (opensearch-project#4541)
Signed-off-by: Lantao Jin <ltjin@amazon.com>
commit 42a415f
Author: qianheng <qianheng@amazon.com>
Date: Tue Oct 14 17:42:45 2025 +0800
Including metadata fields type when doing agg/filter script push down (opensearch-project#4522)
* Including metadata fields type when doing agg/filter script push down
Signed-off-by: Heng Qian <qianheng@amazon.com>
* Fix IT
Signed-off-by: Heng Qian <qianheng@amazon.com>
---------
Signed-off-by: Heng Qian <qianheng@amazon.com>
commit 8de0386
Author: Xinyuan Lu <xinyual@amazon.com>
Date: Tue Oct 14 16:41:08 2025 +0800
Fix percentile bug (opensearch-project#4539)
* fix percentile bug
Signed-off-by: xinyual <xinyual@amazon.com>
* add IT
Signed-off-by: xinyual <xinyual@amazon.com>
* optimize it
Signed-off-by: xinyual <xinyual@amazon.com>
---------
Signed-off-by: xinyual <xinyual@amazon.com>
commit de2fdc8
Author: Lantao Jin <ltjin@amazon.com>
Date: Tue Oct 14 12:29:03 2025 +0800
[FollowUp] Set 0 and negative value of subsearch.maxout as unlimited (opensearch-project#4534)
* [FollowUp] Set 0 and negative value of subsearch.maxout as unlimited
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* fix doctest
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* Fix conflicts
Signed-off-by: Lantao Jin <ltjin@amazon.com>
---------
Signed-off-by: Lantao Jin <ltjin@amazon.com>
commit 977b7ab
Author: Simeon Widdis <sawiddis@gmail.com>
Date: Mon Oct 13 20:23:10 2025 -0700
Update stalled action (opensearch-project#4485)
commit fddbb70
Author: Lantao Jin <ltjin@amazon.com>
Date: Tue Oct 14 10:23:12 2025 +0800
Add configurable sytem limitations for `subsearch` and `join` command (opensearch-project#4501)
* Add configurable sytem limitations for subsearch and join command
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* Fix IT
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* typo
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* fix IT
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* remove rollback in doc
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* address comments
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* fix typo
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* Fix IT
Signed-off-by: Lantao Jin <ltjin@amazon.com>
---------
Signed-off-by: Lantao Jin <ltjin@amazon.com>
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancementNew feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@LantaoJin@anasalkouz@penghuo@Swiddis
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Use _doc + _shard_doc as sort tiebreaker to get better performance - #4569

Merged
penghuo merged 2 commits into
opensearch-project:mainfrom
LantaoJin:pr/sort_tiebreaker
Oct 15, 2025
Merged

Use _doc + _shard_doc as sort tiebreaker to get better performance#4569
penghuo merged 2 commits into
opensearch-project:mainfrom
LantaoJin:pr/sort_tiebreaker

Conversation

@LantaoJin

@LantaoJinLantaoJin commented Oct 15, 2025

Copy link
Copy Markdown
Member

Description

Before #4378, the sort in PIT search is
case 1: if no sort field specified, sort by _doc + _id (+ means "then"). (❎ could cause high memory issue)
case 2: if sort fields specified, sort by fields. (❎ paged results could miss or duplicate hits)
case 3: if sort fields specified and query contains a filter, sort by _doc. (❎ paged results could miss or duplicate hits)

#4378 added the _shard_doc as sort tiebreaker with
case 1: if no sort field specified, sort by _shard_doc. (❎ performance regression)
case 2: if sort fields specified, sort by fields + _shard_doc.(❎ lower performance on low cardinality field)

#4435 found performance regression in case 1 and partially revert the changes to
case 1: if no sort field specified, sort by _doc + _id. (❎ could cause high memory issue)
case 2: if sort fields specified, sort by fields. (❎ paged results could miss or duplicate hits)

After this PR, we change the sort in PIT search to
case 1: if no sort field specified, sort by _doc + _shard_doc. ✅
case 2: if sort fields specified, sort by fields + _doc + _shard_doc.✅

RCA of performance regression:
_shard_doc is not a stored field in index which will be generated in runtime when comparison. Computing _shard_doc per document is a high cost operation. But sorting by _doc then _shard_doc only generates _shard_doc when the _doc values are conflicted.
Even in the case of user specified sort fields, we should sort by fields then _doc then _shard_doc to reduce the computing of _shard_doc. For example, if the sort field is a low cardinality field, e.g. gender, sorting by gender then _doc then _shard_doc generates _shard_doc for comparison only if values of gender and _doc are both conflicted.

This PR is no needed to backport to 2.19-dev since shard_doc feature is only available since OS 3.3.0

Related Issues

Resolves #[Issue number to be closed when this PR is merged]

Check List

  • New functionality includes testing.
  • New functionality has been documented.
  • New functionality has javadoc added.
  • New functionality has a user manual doc added.
  • New PPL command checklist all confirmed.
  • API changes companion pull request created.
  • Commits are signed per the DCO using --signoff or -s.
  • Public documentation issue/PR created.

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.
For more information on following Developer Certificate of Origin and signing off your commits, please check here.

Signed-off-by: Lantao Jin <ltjin@amazon.com>
@LantaoJinLantaoJin added the enhancement New feature or request label Oct 15, 2025
@LantaoJinLantaoJin changed the title Use _shard_doc as sort tiebreaker to get better performanceUse _doc + _shard_doc as sort tiebreaker to get better performanceOct 15, 2025
Signed-off-by: Lantao Jin <ltjin@amazon.com>
// Workaround to preserve sort location more exactly,
// see https://github.com/opensearch-project/sql/pull/3061
this.sourceBuilder.sort(METADATA_FIELD_ID, ASC);
this.sourceBuilder.sort(SortBuilders.shardDocSort());

@SwiddisSwiddisOct 15, 2025

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Does it matter if we duplicate fields in the sorting list? We could simplify/remove the below else logic by just always appending this, I would expect Lucene to optimize it in the background but I haven't measured it.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

not sure will Lucene optimize duplicated fields or _doc in sorting, but for sure the duplicated _shard_doc is not allowed in OpenSearch Core. It is no harmful for restricted checker here.

@anasalkouz

Copy link
Copy Markdown
Member

Can you share the performance benchmark for the 3 approaches?

@LantaoJin

LantaoJin commented Oct 15, 2025

Copy link
Copy Markdown
MemberAuthor

Can you share the performance benchmark for the 3 approaches?

I haven't run the benchmark, the RCA was made by reading the code of Luence and OS shard_doc feature.

The performance of _doc then _shard_doc is same as _doc then _id, provided by @ahkcs on Oct 1st. (the case 1)

For case 2, the current fields + _doc + _shard_doc is an further optimization upon fields + _shard_doc based on above benchmark result with inference.

Will rerun some benchmark to double confirm.

@penghuo
penghuo merged commit 3388dc7 into opensearch-project:mainOct 15, 2025
38 checks passed
ykmr1224 added a commit to ykmr1224/sql that referenced this pull request Oct 15, 2025
commit cba8d02
Author: Tomoyuki MORITA <moritato@amazon.com>
Date: Wed Oct 15 13:08:05 2025 -0700
Add MAP_APPEND internal function to Calcite PPL (opensearch-project#4515)
* Add MAP_APPEND internal function to Calcite PPL
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
* Minor fix
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
* Address comment
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
* Rebase and fix IT issue
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
---------
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
commit 3388dc7
Author: Lantao Jin <ltjin@amazon.com>
Date: Thu Oct 16 01:45:29 2025 +0800
Use `_doc` + `_shard_doc` as sort tiebreaker to get better performance (opensearch-project#4569)
* Use _shard_doc as sort tiebreaker
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* _doc as a part of tie-breaker have better performance
Signed-off-by: Lantao Jin <ltjin@amazon.com>
---------
Signed-off-by: Lantao Jin <ltjin@amazon.com>
commit 5630119
Author: qianheng <qianheng@amazon.com>
Date: Wed Oct 15 16:40:41 2025 +0800
Fix sort push down into agg after project already pushed (opensearch-project#4546)
* Fix sort push down into agg
Signed-off-by: Heng Qian <qianheng@amazon.com>
* Change some json files to yaml format
Signed-off-by: Heng Qian <qianheng@amazon.com>
---------
Signed-off-by: Heng Qian <qianheng@amazon.com>
commit 1e62fba
Author: Tomoyuki MORITA <moritato@amazon.com>
Date: Tue Oct 14 17:20:38 2025 -0700
Fix JsonExtractAllFunctionIT failure (opensearch-project#4556)
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
commit 02ee33e
Author: Kai Huang <105710027+ahkcs@users.noreply.github.com>
Date: Tue Oct 14 14:28:53 2025 -0700
Add more examples to the `where` command doc (opensearch-project#4457)
Co-authored-by: Manasvini B S <manasvis@amazon.com>
commit 0b7e86c
Author: Jialiang Liang <jiallian@amazon.com>
Date: Tue Oct 14 10:46:01 2025 -0700
[Enhancement] Error handling for illegal character usage in java regex named capture group (opensearch-project#4434)
Co-authored-by: Simeon Widdis <sawiddis@amazon.com>
commit 9c97cfb
Author: Tomoyuki MORITA <moritato@amazon.com>
Date: Tue Oct 14 08:36:43 2025 -0700
Add JSON_EXTRACT_ALL internal function for Calcite PPL (opensearch-project#4489)
* Add JSON_EXTRACT_ALL internal function for Calcite PPL
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
* Address comments
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
* Minor fix
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
---------
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
commit 89dbc31
Author: Lantao Jin <ltjin@amazon.com>
Date: Tue Oct 14 18:24:52 2025 +0800
Check server status before starting Prometheus (opensearch-project#4537)
* Check server status before starting Prometheus
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* Change to func call
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* Fix doc
Signed-off-by: Lantao Jin <ltjin@amazon.com>
---------
Signed-off-by: Lantao Jin <ltjin@amazon.com>
commit fe62472
Author: Lantao Jin <ltjin@amazon.com>
Date: Tue Oct 14 18:10:27 2025 +0800
Update request builder after pushdown sort into agg buckets (opensearch-project#4541)
Signed-off-by: Lantao Jin <ltjin@amazon.com>
commit 42a415f
Author: qianheng <qianheng@amazon.com>
Date: Tue Oct 14 17:42:45 2025 +0800
Including metadata fields type when doing agg/filter script push down (opensearch-project#4522)
* Including metadata fields type when doing agg/filter script push down
Signed-off-by: Heng Qian <qianheng@amazon.com>
* Fix IT
Signed-off-by: Heng Qian <qianheng@amazon.com>
---------
Signed-off-by: Heng Qian <qianheng@amazon.com>
commit 8de0386
Author: Xinyuan Lu <xinyual@amazon.com>
Date: Tue Oct 14 16:41:08 2025 +0800
Fix percentile bug (opensearch-project#4539)
* fix percentile bug
Signed-off-by: xinyual <xinyual@amazon.com>
* add IT
Signed-off-by: xinyual <xinyual@amazon.com>
* optimize it
Signed-off-by: xinyual <xinyual@amazon.com>
---------
Signed-off-by: xinyual <xinyual@amazon.com>
commit de2fdc8
Author: Lantao Jin <ltjin@amazon.com>
Date: Tue Oct 14 12:29:03 2025 +0800
[FollowUp] Set 0 and negative value of subsearch.maxout as unlimited (opensearch-project#4534)
* [FollowUp] Set 0 and negative value of subsearch.maxout as unlimited
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* fix doctest
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* Fix conflicts
Signed-off-by: Lantao Jin <ltjin@amazon.com>
---------
Signed-off-by: Lantao Jin <ltjin@amazon.com>
commit 977b7ab
Author: Simeon Widdis <sawiddis@gmail.com>
Date: Mon Oct 13 20:23:10 2025 -0700
Update stalled action (opensearch-project#4485)
commit fddbb70
Author: Lantao Jin <ltjin@amazon.com>
Date: Tue Oct 14 10:23:12 2025 +0800
Add configurable sytem limitations for `subsearch` and `join` command (opensearch-project#4501)
* Add configurable sytem limitations for subsearch and join command
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* Fix IT
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* typo
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* fix IT
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* remove rollback in doc
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* address comments
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* fix typo
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* Fix IT
Signed-off-by: Lantao Jin <ltjin@amazon.com>
---------
Signed-off-by: Lantao Jin <ltjin@amazon.com>
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancementNew feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@LantaoJin@anasalkouz@penghuo@Swiddis
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Use _doc + _shard_doc as sort tiebreaker to get better performance - #4569

Merged
penghuo merged 2 commits into
opensearch-project:mainfrom
LantaoJin:pr/sort_tiebreaker
Oct 15, 2025
Merged

Use _doc + _shard_doc as sort tiebreaker to get better performance#4569
penghuo merged 2 commits into
opensearch-project:mainfrom
LantaoJin:pr/sort_tiebreaker

Conversation

@LantaoJin

@LantaoJinLantaoJin commented Oct 15, 2025

Copy link
Copy Markdown
Member

Description

Before #4378, the sort in PIT search is
case 1: if no sort field specified, sort by _doc + _id (+ means "then"). (❎ could cause high memory issue)
case 2: if sort fields specified, sort by fields. (❎ paged results could miss or duplicate hits)
case 3: if sort fields specified and query contains a filter, sort by _doc. (❎ paged results could miss or duplicate hits)

#4378 added the _shard_doc as sort tiebreaker with
case 1: if no sort field specified, sort by _shard_doc. (❎ performance regression)
case 2: if sort fields specified, sort by fields + _shard_doc.(❎ lower performance on low cardinality field)

#4435 found performance regression in case 1 and partially revert the changes to
case 1: if no sort field specified, sort by _doc + _id. (❎ could cause high memory issue)
case 2: if sort fields specified, sort by fields. (❎ paged results could miss or duplicate hits)

After this PR, we change the sort in PIT search to
case 1: if no sort field specified, sort by _doc + _shard_doc. ✅
case 2: if sort fields specified, sort by fields + _doc + _shard_doc.✅

RCA of performance regression:
_shard_doc is not a stored field in index which will be generated in runtime when comparison. Computing _shard_doc per document is a high cost operation. But sorting by _doc then _shard_doc only generates _shard_doc when the _doc values are conflicted.
Even in the case of user specified sort fields, we should sort by fields then _doc then _shard_doc to reduce the computing of _shard_doc. For example, if the sort field is a low cardinality field, e.g. gender, sorting by gender then _doc then _shard_doc generates _shard_doc for comparison only if values of gender and _doc are both conflicted.

This PR is no needed to backport to 2.19-dev since shard_doc feature is only available since OS 3.3.0

Related Issues

Resolves #[Issue number to be closed when this PR is merged]

Check List

  • New functionality includes testing.
  • New functionality has been documented.
  • New functionality has javadoc added.
  • New functionality has a user manual doc added.
  • New PPL command checklist all confirmed.
  • API changes companion pull request created.
  • Commits are signed per the DCO using --signoff or -s.
  • Public documentation issue/PR created.

By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.
For more information on following Developer Certificate of Origin and signing off your commits, please check here.

Signed-off-by: Lantao Jin <ltjin@amazon.com>
@LantaoJinLantaoJin added the enhancement New feature or request label Oct 15, 2025
@LantaoJinLantaoJin changed the title Use _shard_doc as sort tiebreaker to get better performanceUse _doc + _shard_doc as sort tiebreaker to get better performanceOct 15, 2025
Signed-off-by: Lantao Jin <ltjin@amazon.com>
// Workaround to preserve sort location more exactly,
// see https://github.com/opensearch-project/sql/pull/3061
this.sourceBuilder.sort(METADATA_FIELD_ID, ASC);
this.sourceBuilder.sort(SortBuilders.shardDocSort());

@SwiddisSwiddisOct 15, 2025

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Does it matter if we duplicate fields in the sorting list? We could simplify/remove the below else logic by just always appending this, I would expect Lucene to optimize it in the background but I haven't measured it.

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

not sure will Lucene optimize duplicated fields or _doc in sorting, but for sure the duplicated _shard_doc is not allowed in OpenSearch Core. It is no harmful for restricted checker here.

@anasalkouz

Copy link
Copy Markdown
Member

Can you share the performance benchmark for the 3 approaches?

@LantaoJin

LantaoJin commented Oct 15, 2025

Copy link
Copy Markdown
MemberAuthor

Can you share the performance benchmark for the 3 approaches?

I haven't run the benchmark, the RCA was made by reading the code of Luence and OS shard_doc feature.

The performance of _doc then _shard_doc is same as _doc then _id, provided by @ahkcs on Oct 1st. (the case 1)

For case 2, the current fields + _doc + _shard_doc is an further optimization upon fields + _shard_doc based on above benchmark result with inference.

Will rerun some benchmark to double confirm.

@penghuo
penghuo merged commit 3388dc7 into opensearch-project:mainOct 15, 2025
38 checks passed
ykmr1224 added a commit to ykmr1224/sql that referenced this pull request Oct 15, 2025
commit cba8d02
Author: Tomoyuki MORITA <moritato@amazon.com>
Date: Wed Oct 15 13:08:05 2025 -0700
Add MAP_APPEND internal function to Calcite PPL (opensearch-project#4515)
* Add MAP_APPEND internal function to Calcite PPL
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
* Minor fix
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
* Address comment
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
* Rebase and fix IT issue
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
---------
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
commit 3388dc7
Author: Lantao Jin <ltjin@amazon.com>
Date: Thu Oct 16 01:45:29 2025 +0800
Use `_doc` + `_shard_doc` as sort tiebreaker to get better performance (opensearch-project#4569)
* Use _shard_doc as sort tiebreaker
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* _doc as a part of tie-breaker have better performance
Signed-off-by: Lantao Jin <ltjin@amazon.com>
---------
Signed-off-by: Lantao Jin <ltjin@amazon.com>
commit 5630119
Author: qianheng <qianheng@amazon.com>
Date: Wed Oct 15 16:40:41 2025 +0800
Fix sort push down into agg after project already pushed (opensearch-project#4546)
* Fix sort push down into agg
Signed-off-by: Heng Qian <qianheng@amazon.com>
* Change some json files to yaml format
Signed-off-by: Heng Qian <qianheng@amazon.com>
---------
Signed-off-by: Heng Qian <qianheng@amazon.com>
commit 1e62fba
Author: Tomoyuki MORITA <moritato@amazon.com>
Date: Tue Oct 14 17:20:38 2025 -0700
Fix JsonExtractAllFunctionIT failure (opensearch-project#4556)
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
commit 02ee33e
Author: Kai Huang <105710027+ahkcs@users.noreply.github.com>
Date: Tue Oct 14 14:28:53 2025 -0700
Add more examples to the `where` command doc (opensearch-project#4457)
Co-authored-by: Manasvini B S <manasvis@amazon.com>
commit 0b7e86c
Author: Jialiang Liang <jiallian@amazon.com>
Date: Tue Oct 14 10:46:01 2025 -0700
[Enhancement] Error handling for illegal character usage in java regex named capture group (opensearch-project#4434)
Co-authored-by: Simeon Widdis <sawiddis@amazon.com>
commit 9c97cfb
Author: Tomoyuki MORITA <moritato@amazon.com>
Date: Tue Oct 14 08:36:43 2025 -0700
Add JSON_EXTRACT_ALL internal function for Calcite PPL (opensearch-project#4489)
* Add JSON_EXTRACT_ALL internal function for Calcite PPL
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
* Address comments
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
* Minor fix
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
---------
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
commit 89dbc31
Author: Lantao Jin <ltjin@amazon.com>
Date: Tue Oct 14 18:24:52 2025 +0800
Check server status before starting Prometheus (opensearch-project#4537)
* Check server status before starting Prometheus
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* Change to func call
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* Fix doc
Signed-off-by: Lantao Jin <ltjin@amazon.com>
---------
Signed-off-by: Lantao Jin <ltjin@amazon.com>
commit fe62472
Author: Lantao Jin <ltjin@amazon.com>
Date: Tue Oct 14 18:10:27 2025 +0800
Update request builder after pushdown sort into agg buckets (opensearch-project#4541)
Signed-off-by: Lantao Jin <ltjin@amazon.com>
commit 42a415f
Author: qianheng <qianheng@amazon.com>
Date: Tue Oct 14 17:42:45 2025 +0800
Including metadata fields type when doing agg/filter script push down (opensearch-project#4522)
* Including metadata fields type when doing agg/filter script push down
Signed-off-by: Heng Qian <qianheng@amazon.com>
* Fix IT
Signed-off-by: Heng Qian <qianheng@amazon.com>
---------
Signed-off-by: Heng Qian <qianheng@amazon.com>
commit 8de0386
Author: Xinyuan Lu <xinyual@amazon.com>
Date: Tue Oct 14 16:41:08 2025 +0800
Fix percentile bug (opensearch-project#4539)
* fix percentile bug
Signed-off-by: xinyual <xinyual@amazon.com>
* add IT
Signed-off-by: xinyual <xinyual@amazon.com>
* optimize it
Signed-off-by: xinyual <xinyual@amazon.com>
---------
Signed-off-by: xinyual <xinyual@amazon.com>
commit de2fdc8
Author: Lantao Jin <ltjin@amazon.com>
Date: Tue Oct 14 12:29:03 2025 +0800
[FollowUp] Set 0 and negative value of subsearch.maxout as unlimited (opensearch-project#4534)
* [FollowUp] Set 0 and negative value of subsearch.maxout as unlimited
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* fix doctest
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* Fix conflicts
Signed-off-by: Lantao Jin <ltjin@amazon.com>
---------
Signed-off-by: Lantao Jin <ltjin@amazon.com>
commit 977b7ab
Author: Simeon Widdis <sawiddis@gmail.com>
Date: Mon Oct 13 20:23:10 2025 -0700
Update stalled action (opensearch-project#4485)
commit fddbb70
Author: Lantao Jin <ltjin@amazon.com>
Date: Tue Oct 14 10:23:12 2025 +0800
Add configurable sytem limitations for `subsearch` and `join` command (opensearch-project#4501)
* Add configurable sytem limitations for subsearch and join command
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* Fix IT
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* typo
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* fix IT
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* remove rollback in doc
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* address comments
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* fix typo
Signed-off-by: Lantao Jin <ltjin@amazon.com>
* Fix IT
Signed-off-by: Lantao Jin <ltjin@amazon.com>
---------
Signed-off-by: Lantao Jin <ltjin@amazon.com>
Signed-off-by: Tomoyuki Morita <moritato@amazon.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancementNew feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@LantaoJin@anasalkouz@penghuo@Swiddis