Uh oh!
There was an error while loading. Please reload this page.
[SPARK-19112][CORE] add codec for ZStandard - #17303
Conversation
AmplabJenkins
commented
Mar 15, 2017
Can one of the admins verify this patch? |
srowen
commented
Mar 15, 2017
Same questions from last PR -- can this be something the user includes if needed or is there value in integrating it into Spark? where would it come into play and with what versions of Hadoop et al? |
tgravescs
commented
Mar 15, 2017
this should not be needed just to use to write to hdfs. The regular hadoop input/output type formats have support for it if you are using the right version (I think hadoop 2.8). This seems to be adding the support to the spark.io.compression.codec for internal compression. From what I've heard zstd is better then the other codecs since it gives Gzip level Compression with Lz4 level CPU usage. So if you have a job that had a ton of intermediate data or was causing network issues you may want to use ztsd to get the gzip compression levels without much cpu penalty. @dongjinleekr It doesn't looks like you ran any manual tests on a real cluster? It would be nice to have some basic performance/compression numbers to show it actually working. Are you planning on actually using zstd in your spark deployment? |
rxin
commented
Mar 15, 2017
Yes it'd be nice to have some benchmark on this. |
I did quick benchmarks by using a TPCDS query (Q4) (I just referred the previous work in #10342) |
srowen
commented
May 6, 2017
OK, seems like we should close this. |
| class ZStandardCompressionCodec(conf: SparkConf) extends CompressionCodec { | ||
| override def compressedOutputStream(s: OutputStream): OutputStream = { | ||
| val level = conf.getSizeAsBytes("spark.io.compression.zstandard.level", "3").toInt |
There was a problem hiding this comment.
Use cases which favor speed over size should prefer using level 1.
Compression speed difference can be fairly large.
@Cyan4973 I quickly checked again;
|
Cyan4973
commented
May 9, 2017
@maropu : What about compression ratios ? |
## What changes were proposed in this pull request? This PR proposes to close PRs ... - inactive to the review comments more than a month - WIP and inactive more than a month - with Jenkins build failure but inactive more than a month - suggested to be closed and no comment against that - obviously looking inappropriate (e.g., Branch 0.5) To make sure, I left a comment for each PR about a week ago and I could not have a response back from the author in these PRs below: Closesapache#11129Closesapache#12085Closesapache#12162Closesapache#12419Closesapache#12420Closesapache#12491Closesapache#13762Closesapache#13837Closesapache#13851Closesapache#13881Closesapache#13891Closesapache#13959Closesapache#14091Closesapache#14481Closesapache#14547Closesapache#14557Closesapache#14686Closesapache#15594Closesapache#15652Closesapache#15850Closesapache#15914Closesapache#15918Closesapache#16285Closesapache#16389Closesapache#16652Closesapache#16743Closesapache#16893Closesapache#16975Closesapache#17001Closesapache#17088Closesapache#17119Closesapache#17272Closesapache#17971 Added: Closesapache#17778Closesapache#17303Closesapache#17872 ## How was this patch tested? N/A Author: hyukjinkwon <gurwls223@gmail.com> Closesapache#18017 from HyukjinKwon/close-inactive-prs.
What changes were proposed in this pull request?
Hadoop & HBase started to support ZStandard Compression from their recent releases. This update enables saving a file in HDFS using ZStandard Codec, by implementing ZStandardCodec. It also requires adding a new configuration for default compression level, for example, 'spark.io.compression.zstandard.level.'
How was this patch tested?
3 additional unit tests in
CompressionCodecSuite.scala.