spark-instrumented-optimizer

History

freeman 6c6f325740 [SPARK-5089][PYSPARK][MLLIB] Fix vector convert This is a small change addressing a potentially significant bug in how PySpark + MLlib handles non-float64 numpy arrays. The automatic conversion to `DenseVector` that occurs when passing RDDs to MLlib algorithms in PySpark should automatically upcast to float64s, but currently this wasn't actually happening. As a result, non-float64 would be silently parsed inappropriately during SerDe, yielding erroneous results when running, for example, KMeans. The PR includes the fix, as well as a new test for the correct conversion behavior. davies Author: freeman <the.freeman.lab@gmail.com> Closes #3902 from freeman-lab/fix-vector-convert and squashes the following commits: 764db47 [freeman] Add a test for proper conversion behavior 704f97e [freeman] Return array after changing type		2015-01-05 13:10:59 -08:00
..
docs	[SPARK-4821] [mllib] [python] [docs] Fix for pyspark.mllib.rand doc	2014-12-17 14:12:46 -08:00
lib	[SPARK-2305] [PySpark] Update Py4J to version 0.8.2.1	2014-07-29 19:02:06 -07:00
pyspark	[SPARK-5089][PYSPARK][MLLIB] Fix vector convert	2015-01-05 13:10:59 -08:00
test_support	[SPARK-3634] [PySpark] User's module should take precedence over system modules	2014-09-24 12:10:09 -07:00
.gitignore	[SPARK-3946] gitignore in /python includes wrong directory	2014-10-14 14:09:39 -07:00
run-tests	[SPARK-3721] [PySpark] broadcast objects larger than 2G	2014-11-18 16:17:51 -08:00