spark-instrumented-optimizer

History

Davies Liu 075a0b6582 [SPARK-10917] [SQL] improve performance of complex type in columnar cache This PR improve the performance of complex types in columnar cache by using UnsafeProjection instead of KryoSerializer. A simple benchmark show that this PR could improve the performance of scanning a cached table with complex columns by 15x (comparing to Spark 1.5). Here is the code used to benchmark: ``` df = sc.range(1<<23).map(lambda i: Row(a=Row(b=i, c=str(i)), d=range(10), e=dict(zip(range(10), [str(i) for i in range(10)])))).toDF() df.write.parquet("table") ``` ``` df = sqlContext.read.parquet("table") df.cache() df.count() t = time.time() print df.select("*")._jdf.queryExecution().toRdd().count() print time.time() - t ``` Author: Davies Liu <davies@databricks.com> Closes #8971 from davies/complex.	2015-10-07 15:58:07 -07:00
..
src	[SPARK-10917] [SQL] improve performance of complex type in columnar cache	2015-10-07 15:58:07 -07:00
pom.xml	[SPARK-10300] [BUILD] [TESTS] Add support for test tags in run-tests.py.	2015-10-07 14:11:21 -07:00

Davies Liu 075a0b6582 [SPARK-10917] [SQL] improve performance of complex type in columnar cache

This PR improve the performance of complex types in columnar cache by using UnsafeProjection instead of KryoSerializer.

A simple benchmark show that this PR could improve the performance of scanning a cached table with complex columns by 15x (comparing to Spark 1.5).

Here is the code used to benchmark:

```
df = sc.range(1<<23).map(lambda i: Row(a=Row(b=i, c=str(i)), d=range(10), e=dict(zip(range(10), [str(i) for i in range(10)])))).toDF()
df.write.parquet("table")
```
```
df = sqlContext.read.parquet("table")
df.cache()
df.count()
t = time.time()
print df.select("*")._jdf.queryExecution().toRdd().count()
print time.time() - t
```

Author: Davies Liu <davies@databricks.com>

Closes #8971 from davies/complex.

2015-10-07 15:58:07 -07:00

src

[SPARK-10917] [SQL] improve performance of complex type in columnar cache

2015-10-07 15:58:07 -07:00

pom.xml

[SPARK-10300] [BUILD] [TESTS] Add support for test tags in run-tests.py.

2015-10-07 14:11:21 -07:00