Uh oh!
There was an error while loading. Please reload this page.
[SPARK-20797][MLLIB]fix LocalLDAModel.save() bug. - #18034
Conversation
| // (floatSize * vectorSize + 15) * numWords | ||
| val approxSize = (4L * k + 15) * topicsMatrix.numRows | ||
| val nPartitions = ((approxSize / bufferSize) + 1).toInt | ||
| spark.createDataFrame(topics).repartition(nPartitions).write.parquet(Loader.dataPath(path)) |
There was a problem hiding this comment.
The problem is that this writes multiple files. I don't think you can do that.
There was a problem hiding this comment.
why not? i think it does works. the multiple parquet files will may be in random order, but it will save topic indices. when u call load process, parquet will restore dataframe, u can check the LocalLDAModel's load method, it will scan all dataframe's row with the topic indices to rebuild the (topic x vocab) breeze matrix.
There was a problem hiding this comment.
I believe the point of repartition(1) is to export this to a single file for external consumption. Of course writing it succeeds, but I don't think this is what it is intended to do.
There was a problem hiding this comment.
I've read mllib's online lda's code implements and the paper.
i found a simple way to test it , u can try this code snippet, simulation a matrix to save:
val vocabSize = 500000val k = 300val random = new Random()val topicsDenseMatrix = DenseMatrix.fill[Double](vocabSize, k)(random.nextDouble())
There was a problem hiding this comment.
I think the PR is trying to do something similar to https://github.com/apache/spark/pull/9989/files.
There was a problem hiding this comment.
@srowen what's your concern with saving multiple files instead of one?
We've encountered crippling errors saving large LDA models in the ml api as well. The problem there is even worse, since the entire matrix is treated as one single datum
There was a problem hiding this comment.
Can you add a testcase to save model into multiple files and load back and check the correctness ?
AmplabJenkins
commented
Jun 9, 2018
Can one of the admins verify this patch? |
HyukjinKwon
commented
Jul 16, 2018
Any update? seems the author looks inactive. Please let me know if I am not mistaken. Let me leave this closed if so. |
Closesapache#17422Closesapache#17619Closesapache#18034Closesapache#18229Closesapache#18268Closesapache#17973Closesapache#18125Closesapache#18918Closesapache#19274Closesapache#19456Closesapache#19510Closesapache#19420Closesapache#20090Closesapache#20177Closesapache#20304Closesapache#20319Closesapache#20543Closesapache#20437Closesapache#21261Closesapache#21726Closesapache#14653Closesapache#13143Closesapache#17894Closesapache#19758Closesapache#12951Closesapache#17092Closesapache#21240Closesapache#16910Closesapache#12904Closesapache#21731Closesapache#21095 Added: Closesapache#19233Closesapache#20100Closesapache#21453Closesapache#21455Closesapache#18477 Added: Closesapache#21812Closesapache#21787 Author: hyukjinkwon <gurwls223@apache.org> Closesapache#21781 from HyukjinKwon/closing-prs.
What changes were proposed in this pull request?
LocalLDAModel's model save function has a bug:
please see: https://issues.apache.org/jira/browse/SPARK-20797
add some code like word2vec's save method (repartition), to avoid this bug.
How was this patch tested?
it's hard to test. need large data to train with online LDA method.