[HUDI-3378] RFC-46 rebase - #7345
Merged
Merged
Conversation
wzx140
force-pushed
the
rfc46-rebase
branch
5 times, most recently
from
December 3, 2022 04:14
00cb133 to
7a65409
Compare
Contributor
Author
|
@hudi-bot run azure |
Contributor
Author
alexeykudinkin
force-pushed
the
rfc46-rebase
branch
from
December 13, 2022 06:30
f559ebd to
c09d56e
Compare
alexeykudinkin
self-requested a review
December 13, 2022 06:31
wzx140
force-pushed
the
rfc46-rebase
branch
2 times, most recently
from
December 13, 2022 06:52
2cfcca5 to
1930cfe
Compare
Contributor
Author
|
@hudi-bot run azure |
alexeykudinkin
approved these changes
Dec 13, 2022
alexeykudinkin
left a comment
Contributor
There was a problem hiding this comment.
Great job @wzx140!
alexeykudinkin
force-pushed
the
rfc46-rebase
branch
3 times, most recently
from
December 14, 2022 02:04
1930cfe to
060d8e2
Compare
[minor] add more test for rfc46 (apache#7003) ## Change Logs - Add HoodieSparkValidateDuplicateKeyRecordMerger behaving the same as ValidateDuplicateKeyPayload. We should use it with config "hoodie.sql.insert.mode=strict". - Fix nest field exist in HoodieCatalystExpressionUtils - Fix rewrite in HoodieInternalRowUtiles to support type promoted as avro - Fallback to avro when use "merge into" sql - Fix some schema handling issue - Support delta streamer - Convert parquet schema to spark schema and then avro schema(in org.apache.hudi.io.storage.HoodieSparkParquetReader#getSchema). Some types in avro are not compatible with parquet. For ex, decimal as int32/int64 in parquet will convert to int/long in avro. Because avro do not has decimal as int/long . We will lose the logic type info if we directly convert it to avro schema. - Support schema evolution in parquet block [Minor] fix multi deser avro payload (apache#7021) In HoodieAvroRecord, we will call isDelete, shouldIgnore before we write it to the file. Each method will deserialize HoodiePayload. So we add deserialization method in HoodieRecord and call this method once before calling isDelete or shouldIgnore. Co-authored-by: wangzixuan.wzxuan <wangzixuan.wzxuan@bytedance.com> Co-authored-by: Alexey Kudinkin <alexey@infinilake.com> Co-authored-by: Alexey Kudinkin <alexey.kudinkin@gmail.com> [MINOR] Properly registering target classes w/ Kryo (apache#7026) * Added `HoodieKryoRegistrar` registering necessary Hudi's classes w/ Kryo to make their serialization more efficient (by serializing just the class id, in-liue the fully qualified class-name) * Redirected Kryo registration to `HoodieKryoRegistrar` * Registered additional classes likely to be serialized by Kryo * Updated tests * Fixed serialization of Avro's `Utf8` to serialize just the bytes * Added tests * Added custom `AvroUtf8Serializer`; Tidying up * Extracted `HoodieCommonKryoRegistrar` to leverage in `SerializationUtils` * `HoodieKryoRegistrar` > `HoodieSparkKryoRegistrar`; Rebased `HoodieSparkKryoRegistrar` onto `HoodieCommonKryoRegistrar` * `lint` * Fixing compilation for Spark 2.x * Disabling flaky test [MINOR] Make sure all `HoodieRecord`s are appropriately serializable by Kryo (apache#6977) * Make sure `HoodieRecord`, `HoodieKey`, `HoodieRecordLocation` are all `KryoSerializable` * Revisited `HoodieRecord` serialization hooks to make sure they a) could not be overridden, b) provide for hooks to properly serialize record's payload; Implemented serialization hooks for `HoodieAvroIndexedRecord`; Implemented serialization hooks for `HoodieEmptyRecord`; * Revisited `HoodieRecord` serialization hooks to make sure they a) could not be overridden, b) provide for hooks to properly serialize record's payload; Implemented serialization hooks for `HoodieAvroIndexedRecord`; Implemented serialization hooks for `HoodieEmptyRecord`; Implemented serialization hooks for `HoodieAvroRecord`; * Revisited `HoodieSparkRecord` to transiently hold on to the schema so that it could project row * Implemented serialization hooks for `HoodieSparkRecord` * Added `TestHoodieSparkRecord` * Added tests for Avro-based records * Added test for `HoodieEmptyRecord` * Fixed sealing/unsealing for `HoodieRecord` in `HoodieBackedTableMetadataWriter` * Properly handle deflated records * Fixing `Row`s encoding * Fixed `HoodieRecord` to be properly sealed/unsealed * Fixed serialization of the `HoodieRecordGlobalLocation` [MINOR] Additional fixes for apache#6745 (apache#6947) * Tidying up * Tidying up more * Cleaning up duplication * Tidying up * Revisited legacy operating mode configuration * Tidying up * Cleaned up `projectUnsafe` API * Fixing compilation * Cleaning up `HoodieSparkRecord` ctors; Revisited mandatory unsafe-projection * Fixing compilation * Cleaned up `ParquetReader` initialization * Revisited `HoodieSparkRecord` to accept either `UnsafeRow` or `HoodieInternalRow`, and avoid unnecessary copying after unsafe-projection * Cleaning up redundant exception spec * Make sure `updateMetadataFields` properly wraps `InternalRow` into `HoodieInternalRow` if necessary; Cleaned up `MetadataValues` * Fixed meta-fields extraction and `HoodieInternalRow` composition w/in `HoodieSparkRecord` * De-duplicate `HoodieSparkRecord` ctors; Make sure either only `UnsafeRow` or `HoodieInternalRow` are permitted inside `HoodieSparkRecord` * Removed unnecessary copying * Cleaned up projection for `HoodieSparkRecord` (dropping partition columns); Removed unnecessary copying * Fixing compilation * Fixing compilation (for Flink) * Cleaned up File Raders' interfaces: - Extracted `HoodieSeekingFileReader` interface (for key-ranged reads) - Pushed down concrete implementation methods into `HoodieAvroFileReaderBase` from the interfaces * Cleaned up File Readers impls (inline with then new interfaces) * Rebsaed `HoodieBackedTableMetadata` onto new `HoodieSeekingFileReader` * Tidying up * Missing licenses * Re-instate custom override for `HoodieAvroParquetReader`; Tidying up * Fixed missing cloning w/in `HoodieLazyInsertIterable` * Fixed missing cloning in deduplication flow * Allow `HoodieSparkRecord` to hold `ColumnarBatchRow` * Missing licenses * Fixing compilation * Missing changes * Fixed Spark 2.x validation whether the row was read as a batch Fix comment in RFC46 (apache#6745) * rename * add MetadataValues in updateMetadataValues * remove singleton in fileFactory * add truncateRecordKey * remove hoodieRecord#setData * rename HoodieAvroRecord * fix code style * fix HoodieSparkRecordSerializer * fix benchmark * fix SparkRecordUtils * instantiate HoodieWriteConfig on the fly * add test * fix HoodieSparkRecordSerializer. Replace Java's object serialization with kryo * add broadcast * fix comment * remove unnecessary broadcast * add unsafe check in spark record * fix getRecordColumnValues * remove spark.sql.parquet.writeLegacyFormat * fix unsafe projection * fix * pass external schema * update doc * rename back to HoodieAvroRecord * fix * remove comparable wrapper * fix comment * fix comment * fix comment * fix comment * simplify row copy * fix ParquetReaderIterator Co-authored-by: Shawy Geng <gengxiaoyu1996@gmail.com> Co-authored-by: wangzixuan.wzxuan <wangzixuan.wzxuan@bytedance.com> [RFC-46][HUDI-4414] Update the RFC-46 doc to fix comments feedback (apache#6132) * Update the RFC-46 doc to fix comments feedback * fix Co-authored-by: wangzixuan.wzxuan <wangzixuan.wzxuan@bytedance.com> [HUDI-4301] [HUDI-3384][HUDI-3385] Spark specific file reader/writer.(apache#5629) * [HUDI-4301] [HUDI-3384][HUDI-3385] Spark specific file reader/writer. * add schema finger print * add benchmark * a new way to config the merger * fix Co-authored-by: wangzixuan.wzxuan <wangzixuan.wzxuan@bytedance.com> Co-authored-by: gengxiaoyu <gengxiaoyu@bytedance.com> [HUDI-3350][HUDI-3351] Support HoodieMerge API and Spark engine-specific HoodieRecord (apache#5627) Co-authored-by: wangzixuan.wzxuan <wangzixuan.wzxuan@bytedance.com> [HUDI-4344] fix usage of HoodieDataBlock#getRecordIterator (apache#6005) Co-authored-by: wangzixuan.wzxuan <wangzixuan.wzxuan@bytedance.com> [HUDI-4292][RFC-46] Update doc to align with the Record Merge API changes (apache#5927) [MINOR] Fix type casting in TestHoodieHFileReaderWriter [HUDI-3378][HUDI-3379][HUDI-3381] Migrate usage of HoodieRecordPayload and raw Avro payload to HoodieRecord (apache#5522) Co-authored-by: Alexey Kudinkin <alexey@infinilake.com> Co-authored-by: wangzixuan.wzxuan <wangzixuan.wzxuan@bytedance.com>
…o multi record keys.
Contributor
Author
|
@hudi-bot run azure |
1 similar comment
Contributor
Author
|
@hudi-bot run azure |
Contributor
Author
Member
Member
|
great job! kudos to @wzx140 @alexeykudinkin @minihippo ! |
4 tasks
danny0405
added a commit
to danny0405/hudi
that referenced
this pull request
Mar 15, 2023
4 tasks
danny0405
added a commit
that referenced
this pull request
Mar 16, 2023
danny0405
added a commit
to danny0405/hudi
that referenced
this pull request
Mar 23, 2023
…ache#8192) Introduced by apache#7345 (cherry picked from commit eb92159)
nsivabalan
pushed a commit
to nsivabalan/hudi
that referenced
this pull request
Mar 23, 2023
…ache#8192) Introduced by apache#7345 (cherry picked from commit eb92159)
fengjian428
pushed a commit
to fengjian428/hudi
that referenced
this pull request
Apr 5, 2023
## Change Logs - Add HoodieSparkValidateDuplicateKeyRecordMerger behaving the same as ValidateDuplicateKeyPayload. We should use it with config "hoodie.sql.insert.mode=strict". - Fix nest field exist in HoodieCatalystExpressionUtils - Fix rewrite in HoodieInternalRowUtiles to support type promoted as avro - Fallback to avro when use "merge into" sql - Fix some schema handling issue - Support delta streamer - Convert parquet schema to spark schema and then avro schema(in org.apache.hudi.io.storage.HoodieSparkParquetReader#getSchema). Some types in avro are not compatible with parquet. For ex, decimal as int32/int64 in parquet will convert to int/long in avro. Because avro do not has decimal as int/long . We will lose the logic type info if we directly convert it to avro schema. - Support schema evolution in parquet block [Minor] fix multi deser avro payload (apache#7021) In HoodieAvroRecord, we will call isDelete, shouldIgnore before we write it to the file. Each method will deserialize HoodiePayload. So we add deserialization method in HoodieRecord and call this method once before calling isDelete or shouldIgnore. Co-authored-by: wangzixuan.wzxuan <wangzixuan.wzxuan@bytedance.com> Co-authored-by: Alexey Kudinkin <alexey@infinilake.com> Co-authored-by: Alexey Kudinkin <alexey.kudinkin@gmail.com> [MINOR] Properly registering target classes w/ Kryo (apache#7026) * Added `HoodieKryoRegistrar` registering necessary Hudi's classes w/ Kryo to make their serialization more efficient (by serializing just the class id, in-liue the fully qualified class-name) [MINOR] Make sure all `HoodieRecord`s are appropriately serializable by Kryo (apache#6977) * Make sure `HoodieRecord`, `HoodieKey`, `HoodieRecordLocation` are all `KryoSerializable` * Revisited `HoodieRecord` serialization hooks to make sure they a) could not be overridden, b) provide for hooks to properly serialize record's payload; Implemented serialization hooks for `HoodieAvroIndexedRecord`; Implemented serialization hooks for `HoodieEmptyRecord`; Implemented serialization hooks for `HoodieAvroRecord`; * Revisited `HoodieSparkRecord` to transiently hold on to the schema so that it could project row [MINOR] Additional fixes for apache#6745 (apache#6947) Co-authored-by: Shawy Geng <gengxiaoyu1996@gmail.com> Co-authored-by: wangzixuan.wzxuan <wangzixuan.wzxuan@bytedance.com> [RFC-46][HUDI-4414] Update the RFC-46 doc to fix comments feedback (apache#6132) Co-authored-by: wangzixuan.wzxuan <wangzixuan.wzxuan@bytedance.com> [HUDI-4301] [HUDI-3384][HUDI-3385] Spark specific file reader/writer.(apache#5629) Co-authored-by: wangzixuan.wzxuan <wangzixuan.wzxuan@bytedance.com> Co-authored-by: gengxiaoyu <gengxiaoyu@bytedance.com> [HUDI-3350][HUDI-3351] Support HoodieMerge API and Spark engine-specific HoodieRecord (apache#5627) Co-authored-by: wangzixuan.wzxuan <wangzixuan.wzxuan@bytedance.com> [HUDI-4344] fix usage of HoodieDataBlock#getRecordIterator (apache#6005) Co-authored-by: wangzixuan.wzxuan <wangzixuan.wzxuan@bytedance.com> [HUDI-4292][RFC-46] Update doc to align with the Record Merge API changes (apache#5927) [MINOR] Fix type casting in TestHoodieHFileReaderWriter [HUDI-3378][HUDI-3379][HUDI-3381] Migrate usage of HoodieRecordPayload and raw Avro payload to HoodieRecord (apache#5522) Co-authored-by: Alexey Kudinkin <alexey@infinilake.com> Co-authored-by: wangzixuan.wzxuan <wangzixuan.wzxuan@bytedance.com> Co-authored-by: gengxiaoyu <gengxiaoyu@bytedance.com>
fengjian428
pushed a commit
to fengjian428/hudi
that referenced
this pull request
Apr 5, 2023
stayrascal
pushed a commit
to stayrascal/hudi
that referenced
this pull request
Apr 20, 2023
h1ap
pushed a commit
to h1ap/hudi
that referenced
this pull request
May 15, 2023
4 tasks
Contributor
|
This PR introduces so many regressions, don't think it is a valuable PR from the gains. Let's be more conservertive to accept huge PRs and more rigorous for code reviewing without sufficient tests. |
4 tasks
KnightChess
pushed a commit
to KnightChess/hudi
that referenced
this pull request
Jan 2, 2024
…ache#8192) Introduced by apache#7345 (cherry picked from commit eb92159)
ThinkerLei
reviewed
Feb 22, 2024
| @SuppressWarnings("unchecked") | ||
| @Override | ||
| protected final IndexedRecord readRecordPayload(Kryo kryo, Input input) { | ||
| // NOTE: We're leveraging Spark's default [[GenericAvroSerializer]] to serialize Avro |
Contributor
There was a problem hiding this comment.
@wzx140 GenericRecord.class is interface , this will cause InstantiationError. yes or no ?
This was referenced Jun 25, 2026
voonhous
added a commit
to yihua/hudi
that referenced
this pull request
Aug 28, 2026
…ider javadoc names its real caller - HoodieSparkValidateDuplicateKeyRecordMerger: from its introduction in apache#7345 the merger appeared in production sources only as a TODO in ProvidesHoodieConfig, beside the ValidateDuplicateKeyPayload selection that apache#12588 deleted, so the class doc and the test comment now say it was never wired in-repo rather than unwired since apache#12588. - org.apache.hudi.cli.SchemaProvider: the class javadoc claimed Hudi Streamer as its user; it now names the CLI bootstrap path (BootstrapExecutorUtils loading the configured implementation by class name) and points to org.apache.hudi.utilities.schema for the Streamer provider.
3 tasks
voonhous
added a commit
that referenced
this pull request
Aug 28, 2026
…9164) One suite per class, each in the package of the class it covers: - TestCachingIterator (hudi-spark-common, org.apache.hudi.util): full iteration and hasNext idempotency against a counting source iterator. The trait has had no in-repo user since #17457 moved the datasource read paths onto the file-group reader; removing it is left to a separate cleanup. - TestHoodieCatalog: HoodieCatalog.convertTransforms for identity columns, a single bucket transform, a sorted bucket transform keeping its sort columns, and the two rejections (multiple bucket transforms, unsupported transform). HoodieSparkCatalogUtils.MatchBucketTransform is covered through it, including the sorted_bucket arm. - TestBasicStagedTable (added by #19251): commit leaves the catalog alone, abort drops the staged table through it. - TestHoodieSparkValidateDuplicateKeyRecordMerger: the strategy id pinned as the literal, since a table that loads the merger through hoodie.write.record.merge.custom.implementation.classes selects it by that id and persists it, plus the pre-combining fallback to DefaultSparkRecordMerger. The merger was never wired in-repo: from #7345 it appeared in production sources only as a TODO in ProvidesHoodieConfig, beside the ValidateDuplicateKeyPayload selection that #12588 deleted; the class doc now says so instead of pointing at that deleted payload. - TestProcedureParameterImpl (added by #19161): the factories' required/default semantics, toString field presence, a hashCode inequality case, and an equals case that isolates required from default. - TestCliSchemaProvider: the CLI bootstrap SchemaProvider in org.apache.hudi.cli (distinct from the Streamer one; named to avoid the TestSchemaProvider stubs in hudi-utilities and hudi-kafka-connect), shaped like flink's TestSchemaProviderCompatibility: legacy provider, modern provider whose target falls back to the overridden source HoodieSchema, and null handling. The getTargetHoodieSchema catch comment now names the two provider shapes that reach it. Production changes are limited to those two comments; no behavior change. --------- Co-authored-by: voon <voonhousu@gmail.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.

Change Logs
RFC46 Phase 1
To achieve stated goals of avoiding unnecessary conversions into intermediate representation (Avro), we promote HoodieRecord to become a standardized API of interacting with a single record holding internal engine-specific representation of the payload and replace all accesses from HoodieRecordPayload. We extract Record Combining (Merge) API from HoodieRecordPayload into a standalone, stateless component (engine). More information can be found in RFC46.
HoodieAvroRecordMerger and HoodiePayload are used together to support all existing features. Now the spark part has been realized. The following is the caveats with HoodieSparkRecordMerger
Impact
Rebase release-feature-rfc46 on master
Risk level (write none, low medium or high below)
none
Documentation Update
Describe any necessary documentation update if there is any new feature, config, or user-facing change
ticket number here and follow the instruction to make
changes to the website.
Contributor's checklist