Uh oh!
There was an error while loading. Please reload this page.
GH-3530: Optimize DICTIONARY encoding/decoding data structures and use ByteBuffer - #3566
Open
iemejia wants to merge 1 commit into
Open
GH-3530: Optimize DICTIONARY encoding/decoding data structures and use ByteBuffer#3566iemejia wants to merge 1 commit into
iemejia wants to merge 1 commit into
Conversation
This was referenced May 17, 2026
iemejiaforce-pushed
the
parquet-perf-v2-par2-dictionary
branch
from
July 11, 2026 06:15
399c7bb to
a0d51b6Compare…and use ByteBuffer Encoding improvements: - Replace LinkedOpenHashMap with OpenHashMap + ArrayList for all DictionaryValuesWriter subclasses, eliminating insertion-order overhead and enabling O(1) indexed access for dictionary page serialization and fallback - Make IntList.size() O(1) by tracking totalSize incrementally instead of summing across slab arrays Decoding improvements: - Convert PlainValuesDictionary numeric constructors (INT32, INT64, FLOAT, DOUBLE) from InputStream-based per-byte reads to direct ByteBuffer.getInt/getLong/getFloat/getDouble (JVM intrinsics) Binary hashCode caching: - Cache hashCode() for Binary instances that are not backed by reusable byte arrays, avoiding redundant recomputation during dictionary lookups (hash map probes) JMH benchmarks: - DictionaryEncodingBenchmark: scalar encoding for INT32, INT64, FLOAT, DOUBLE, BINARY, and FIXED_LEN_BYTE_ARRAY with LOW/HIGH cardinality and variable-length string/FLBA dimensions - DictionaryDecodingBenchmark: scalar decoding for all types with matching parameterization - TestDataFactory: shared data generation utility for reproducible benchmark inputs - BenchmarkEncodingUtils: helper to drain DictionaryValuesWriter into encoded dictionary page + data bytes for decoder setup
iemejiaforce-pushed
the
parquet-perf-v2-par2-dictionary
branch
from
July 14, 2026 18:39
a0d51b6 to
99962e1Compareiemejia
commented
Aug 11, 2026
MemberAuthor
This one is probably the second one with the most impact of the PRs, in case you have some cycles or know of someone who can take a look at this @Fokko . Notice that the core changes are small, the extras are the benchmark and the tests, but the core PR is smaller than it seems. |
abstractdog added a commit
to abstractdog/parquet-java
that referenced
this pull request
Sep 7, 2026
Replace the hand-rolled scalar byte loops in Binary#equals, Binary#lexicographicCompare, and Binary#hashCode with the JDK 9+ Arrays.equals(byte[],int,int,byte[],int,int), Arrays.compareUnsigned(byte[],int,int,byte[],int,int), Arrays.hashCode, and ByteBuffer.mismatch APIs. These range overloads route through ArraysSupport.vectorizedMismatch, an @IntrinsicCandidate helper HotSpot substitutes with a SIMD byte-scan (SSE / AVX2 / NEON on modern hardware). Measured on this project's BinaryComparisonBenchmark (JDK 17, 1 fork, 3x1s warmup, 5x1s measurement, throughput): equalsMatch_bytesBytes len=512 9.1M -> 39.1M ops/s (4.30x) equalsMatch_bytesBytes len= 64 41.4M -> 160.9M ops/s (3.89x) equalsMismatch_bytesBytes len=512 14.6M -> 55.5M ops/s (3.79x) compareTo_bytesBytes len=512 15.1M -> 58.8M ops/s (3.91x) compareTo_bytesBytes len= 64 59.0M -> 199.4M ops/s (3.38x) equalsMismatch_bytesBuf len= 64 63.8M -> 190.2M ops/s (2.98x) Short values (len=8) see only 1.1-1.7x because there's barely enough work for one SIMD lane. hashCode is separate. The 31*h+b polynomial has a serial dependency across iterations, so SIMD needs a lane-split algebraic trick that HotSpot only gained in JDK 21 (via ArraysSupport.vectorizedHashCode, which is @IntrinsicCandidate). On JDK 17 -- this project's target -- Arrays.hashCode is a plain scalar loop and shows no measured speedup here. The call is kept for forward compatibility: JDK 21+ runtimes pick up the vectorized intrinsic silently; a bespoke loop would stay stuck at scalar. These primitives sit on Binary equality, comparison and hash-map probing, which is the hot code path for statistics min/max maintenance, dictionary probing, predicate evaluation, and bloom-filter build -- i.e. broadly amortized across reader and writer. Semantics are preserved bit-for-bit: - hashCode(byte[]) uses the same 31*h + b polynomial (Arrays.hashCode on full arrays; the polynomial expanded for slices). - equals is bytewise identity. - lexicographicCompare is unsigned bytewise with shorter-first tie-break on prefix match, matching Arrays.compareUnsigned's contract. The change is JIT-friendly across the four Binary/Binary shape pairs (byte[]/byte[], byte[]/ByteBuffer, ByteBuffer/ByteBuffer) and for heap-backed ByteBuffers unwraps to the intrinsic byte[] path via buffer.array(). Follows the same technique already accepted for DeltaByteArrayWriter in PR apache#3465. Complementary to the Binary hashCode caching in the open dictionary-optimization PR apache#3566: caching reduces call frequency; intrinsics speed up the calls that remain plus equals/compareTo which caching does not address. Includes a JMH micro-benchmark (parquet-benchmarks/BinaryComparisonBenchmark) covering equals / compareTo / hashCode across length regimes 8, 64 and 512 for the three Binary shape combinations, with both worst-case (full match) and average-case (mid-length mismatch) inputs.
abstractdog added a commit
to abstractdog/parquet-java
that referenced
this pull request
Sep 7, 2026
Replace the hand-rolled scalar byte loops in Binary#equals, Binary#lexicographicCompare, and Binary#hashCode with the JDK 9+ Arrays.equals(byte[],int,int,byte[],int,int), Arrays.compareUnsigned(byte[],int,int,byte[],int,int), Arrays.hashCode, and ByteBuffer.mismatch APIs. These range overloads route through ArraysSupport.vectorizedMismatch, an @IntrinsicCandidate helper HotSpot substitutes with a SIMD byte-scan (SSE / AVX2 / NEON on modern hardware). Measured on this project's BinaryComparisonBenchmark (JDK 17, 1 fork, 3x1s warmup, 5x1s measurement, throughput): equalsMatch_bytesBytes len=512 9.1M -> 39.1M ops/s (4.30x) equalsMatch_bytesBytes len= 64 41.4M -> 160.9M ops/s (3.89x) equalsMismatch_bytesBytes len=512 14.6M -> 55.5M ops/s (3.79x) compareTo_bytesBytes len=512 15.1M -> 58.8M ops/s (3.91x) compareTo_bytesBytes len= 64 59.0M -> 199.4M ops/s (3.38x) equalsMismatch_bytesBuf len= 64 63.8M -> 190.2M ops/s (2.98x) Short values (len=8) see only 1.1-1.7x because there's barely enough work for one SIMD lane. hashCode is separate. The 31*h+b polynomial has a serial dependency across iterations, so SIMD needs a lane-split algebraic trick that HotSpot only gained in JDK 21 (via ArraysSupport.vectorizedHashCode, which is @IntrinsicCandidate). On JDK 17 -- this project's target -- Arrays.hashCode is a plain scalar loop and shows no measured speedup here. The call is kept for forward compatibility: JDK 21+ runtimes pick up the vectorized intrinsic silently; a bespoke loop would stay stuck at scalar. These primitives sit on Binary equality, comparison and hash-map probing, which is the hot code path for statistics min/max maintenance, dictionary probing, predicate evaluation, and bloom-filter build -- i.e. broadly amortized across reader and writer. Semantics are preserved bit-for-bit: - hashCode(byte[]) uses the same 31*h + b polynomial (Arrays.hashCode on full arrays; the polynomial expanded for slices). - equals is bytewise identity. - lexicographicCompare is unsigned bytewise with shorter-first tie-break on prefix match, matching Arrays.compareUnsigned's contract. The change is JIT-friendly across the four Binary/Binary shape pairs (byte[]/byte[], byte[]/ByteBuffer, ByteBuffer/ByteBuffer) and for heap-backed ByteBuffers unwraps to the intrinsic byte[] path via buffer.array(). Follows the same technique already accepted for DeltaByteArrayWriter in PR apache#3465. Complementary to the Binary hashCode caching in the open dictionary-optimization PR apache#3566: caching reduces call frequency; intrinsics speed up the calls that remain plus equals/compareTo which caching does not address. Includes a JMH micro-benchmark (parquet-benchmarks/BinaryComparisonBenchmark) covering equals / compareTo / hashCode across length regimes 8, 64 and 512 for the three Binary shape combinations, with both worst-case (full match) and average-case (mid-length mismatch) inputs.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for freeto join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Part of #3530 — Apache Parquet Java Performance Improvements
Summary
Optimize dictionary encoding and decoding data structures.
Encoding:
LinkedOpenHashMapwithOpenHashMap+ArrayListfor allDictionaryValuesWritersubclasses, eliminating insertion-order linked-list overhead and enabling O(1) indexed access for dictionary page serialization and fallback.IntList.size()O(1) by trackingtotalSizeincrementally instead of summing across slab arrays.Decoding:
PlainValuesDictionarynumeric constructors (INT32, INT64, FLOAT, DOUBLE) fromInputStream-based per-byte reads to directByteBuffer.getInt/getLong/getFloat/getDouble.Binary hashCode caching:
hashCode()forBinaryinstances not backed by reusable byte arrays, avoiding redundant recomputation during dictionary hash-map probes.JMH benchmarks:
DictionaryEncodingBenchmark,DictionaryDecodingBenchmarkwithTestDataFactoryandBenchmarkEncodingUtils.Benchmark results
Environment: JDK 25.0.3 (Temurin), OpenJDK 64-Bit Server VM, JMH 1.37, Linux x86_64.
Encoding (100K values/iteration, 2 averaged runs):
The extreme Binary LOW_CARD speedup (up to ~100x for len=1000) is due to eliminating
LinkedOpenHashMapper-entry linked-list overhead, autoboxing, andBinary.hashCode()recomputation. With only ~100 distinct values in the hash map, the old code spent most time onhashCode()over the full key bytes at every probe.Decoding: ~1.0x across all types (the
ByteBufferconstructor optimization is once per row group; per-value decode is an array index lookup and was not changed).