spark-instrumented-optimizer

History

Michael Armbrust 6d0633e3ec [SPARK-7548] [SQL] Add explode function for DataFrames Add an `explode` function for dataframes and modify the analyzer so that single table generating functions can be present in a select clause along with other expressions. There are currently the following restrictions: - only top level TGFs are allowed (i.e. no `select(explode('list) + 1)`) - only one may be present in a single select to avoid potentially confusing implicit Cartesian products. TODO: - [ ] Python Author: Michael Armbrust <michael@databricks.com> Closes #6107 from marmbrus/explodeFunction and squashes the following commits: 7ee2c87 [Michael Armbrust] whitespace 6f80ba3 [Michael Armbrust] Update dataframe.py c176c89 [Michael Armbrust] Merge remote-tracking branch 'origin/master' into explodeFunction 81b5da3 [Michael Armbrust] style d3faa05 [Michael Armbrust] fix self join case f9e1e3e [Michael Armbrust] fix python, add since 4f0d0a9 [Michael Armbrust] Merge remote-tracking branch 'origin/master' into explodeFunction e710fe4 [Michael Armbrust] add java and python 52ca0dc [Michael Armbrust] [SPARK-7548][SQL] Add explode function for dataframes.		2015-05-14 19:49:44 -07:00
..
ml	[SPARK-7619] [PYTHON] fix docstring signature	2015-05-14 18:16:22 -07:00
mllib	[SPARK-6092] [MLLIB] Add RankingMetrics in PySpark/MLlib	2015-05-11 09:14:20 -07:00
sql	[SPARK-7548] [SQL] Add explode function for DataFrames	2015-05-14 19:49:44 -07:00
streaming	[SPARK-2808][Streaming][Kafka] update kafka to 0.8.2	2015-05-01 17:54:56 -07:00
__init__.py	[SPARK-4172] [PySpark] Progress API in Python	2015-02-17 13:36:43 -08:00
accumulators.py	[SPARK-6661] Python type errors should print type, not object	2015-04-20 10:44:09 -07:00
broadcast.py	[SPARK-4897] [PySpark] Python 3 support	2015-04-16 16:20:57 -07:00
cloudpickle.py	[SPARK-4897] [PySpark] Python 3 support	2015-04-16 16:20:57 -07:00
conf.py	[SPARK-4897] [PySpark] Python 3 support	2015-04-16 16:20:57 -07:00
context.py	[SPARK-3444] Provide an easy way to change log level	2015-05-01 18:02:51 -07:00
daemon.py	[SPARK-4897] [PySpark] Python 3 support	2015-04-16 16:20:57 -07:00
files.py	[SPARK-3309] [PySpark] Put all public API in __all__	2014-09-03 11:49:45 -07:00
heapq3.py	[SPARK-4897] [PySpark] Python 3 support	2015-04-16 16:20:57 -07:00
java_gateway.py	[SPARK-6949] [SQL] [PySpark] Support Date/Timestamp in Column expression	2015-04-21 00:08:18 -07:00
join.py	[SPARK-4897] [PySpark] Python 3 support	2015-04-16 16:20:57 -07:00
profiler.py	[SPARK-4897] [PySpark] Python 3 support	2015-04-16 16:20:57 -07:00
rdd.py	[SPARK-7438] [SPARK CORE] Fixed validation of relativeSD in countApproxDistinct	2015-05-09 10:03:15 +01:00
rddsampler.py	[SPARK-4897] [PySpark] Python 3 support	2015-04-16 16:20:57 -07:00
resultiterable.py	[SPARK-3074] [PySpark] support groupByKey() with single huge key	2015-04-09 17:07:23 -07:00
serializers.py	[SPARK-4897] [PySpark] Python 3 support	2015-04-16 16:20:57 -07:00
shell.py	[SPARK-4897] [PySpark] Python 3 support	2015-04-16 16:20:57 -07:00
shuffle.py	[SPARK-6953] [PySpark] speed up python tests	2015-04-21 17:49:55 -07:00
statcounter.py	[SPARK-4897] [PySpark] Python 3 support	2015-04-16 16:20:57 -07:00
status.py	[SPARK-4172] [PySpark] Progress API in Python	2015-02-17 13:36:43 -08:00
storagelevel.py	[SPARK-3417] Use new-style classes in PySpark	2014-09-08 15:45:36 -07:00
tests.py	[SPARK-7438] [SPARK CORE] Fixed validation of relativeSD in countApproxDistinct	2015-05-09 10:03:15 +01:00
traceback_utils.py	[SPARK-1087] Move python traceback utilities into new traceback_utils.py file.	2014-09-15 19:28:17 -07:00
worker.py	[SPARK-4897] [PySpark] Python 3 support	2015-04-16 16:20:57 -07:00