1a9c6cddad
Register MLlib's Vector as a SQL user-defined type (UDT) in both Scala and Python. With this PR, we can easily map a RDD[LabeledPoint] to a SchemaRDD, and then select columns or save to a Parquet file. Examples in Scala/Python are attached. The Scala code was copied from jkbradley. ~~This PR contains the changes from #3068 . I will rebase after #3068 is merged.~~ marmbrus jkbradley Author: Xiangrui Meng <meng@databricks.com> Closes #3070 from mengxr/SPARK-3573 and squashes the following commits: 3a0b6e5 [Xiangrui Meng] organize imports 236f0a0 [Xiangrui Meng] register vector as UDT and provide dataset examples |
||
---|---|---|
.. | ||
__init__.py | ||
classification.py | ||
clustering.py | ||
common.py | ||
feature.py | ||
linalg.py | ||
random.py | ||
recommendation.py | ||
regression.py | ||
stat.py | ||
tests.py | ||
tree.py | ||
util.py |