spark-instrumented-optimizer

History

Liquan Pei 3c8fa50590 [SPARK-3097][MLlib] Word2Vec performance improvement mengxr Please review the code. Adding weights in reduceByKey soon. Only output model entry for words appeared in the partition before merging and use reduceByKey to combine model. In general, this implementation is 30s or so faster than implementation using big array. Author: Liquan Pei <liquanpei@gmail.com> Closes #1932 from Ishiihara/Word2Vec-improve2 and squashes the following commits: d5377a9 [Liquan Pei] use syn0Global and syn1Global to represent model cad2011 [Liquan Pei] bug fix for synModify array out of bound 083aa66 [Liquan Pei] update synGlobal in place and reduce synOut size 9075e1c [Liquan Pei] combine syn0Global and syn1Global to synGlobal aa2ab36 [Liquan Pei] use reduceByKey to combine models	2014-08-17 23:29:44 -07:00
..
main/scala/org/apache/spark/mllib	[SPARK-3097][MLlib] Word2Vec performance improvement	2014-08-17 23:29:44 -07:00
test	[SPARK-3087][MLLIB] fix col indexing bug in chi-square and add a check for number of distinct values	2014-08-17 20:53:18 -07:00

Liquan Pei 3c8fa50590 [SPARK-3097][MLlib] Word2Vec performance improvement

mengxr Please review the code. Adding weights in reduceByKey soon.

Only output model entry for words appeared in the partition before merging and use reduceByKey to combine model. In general, this implementation is 30s or so faster than implementation using big array.

Author: Liquan Pei <liquanpei@gmail.com>

Closes #1932 from Ishiihara/Word2Vec-improve2 and squashes the following commits:

d5377a9 [Liquan Pei] use syn0Global and syn1Global to represent model
cad2011 [Liquan Pei] bug fix for synModify array out of bound
083aa66 [Liquan Pei] update synGlobal in place and reduce synOut size
9075e1c [Liquan Pei] combine syn0Global and syn1Global to synGlobal
aa2ab36 [Liquan Pei] use reduceByKey to combine models

2014-08-17 23:29:44 -07:00

main/scala/org/apache/spark/mllib

[SPARK-3097][MLlib] Word2Vec performance improvement

2014-08-17 23:29:44 -07:00

test

[SPARK-3087][MLLIB] fix col indexing bug in chi-square and add a check for number of distinct values

2014-08-17 20:53:18 -07:00