Uh oh!
There was an error while loading. Please reload this page.
[SPARK-26709][SQL] OptimizeMetadataOnlyQuery does not handle empty records correctly - #23635
Closed
gengliangwang wants to merge 7 commits into
Closed
[SPARK-26709][SQL] OptimizeMetadataOnlyQuery does not handle empty records correctly#23635gengliangwang wants to merge 7 commits into
gengliangwang wants to merge 7 commits into
Conversation
Loading
Uh oh!
There was an error while loading. Please reload this page.
Sign up for freeto join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What changes were proposed in this pull request?
When reading from empty tables, the optimization
OptimizeMetadataOnlyQuerymay return wrong results:The result is supposed to be
null. However, with the optimization the result is5.The rule is originally ported from https://issues.apache.org/jira/browse/HIVE-1003 in #13494. In Hive, the rule is disabled by default in a later release(https://issues.apache.org/jira/browse/HIVE-15397), due to the same problem.
It is hard to completely avoid the correctness issue. Because data sources like Parquet can be metadata-only. Spark can't tell whether it is empty or not without actually reading it. This PR disable the optimization by default.
How was this patch tested?
Unit test