History

DB Tsai 681845fd62 [SPARK-24402][SQL] Optimize `In` expression when only one element in the collection or collection is empty ## What changes were proposed in this pull request? Two new rules in the logical plan optimizers are added. 1. When there is only one element in the `Collection`, the physical plan will be optimized to `EqualTo`, so predicate pushdown can be used. ```scala profileDF.filter( $"profileID".isInCollection(Set(6))).explain(true) """ \|== Physical Plan == \|(1) Project [profileID#0] \|+- (1) Filter (isnotnull(profileID#0) && (profileID#0 = 6)) \| +- (1) FileScan parquet [profileID#0] Batched: true, Format: Parquet, \| PartitionFilters: [], \| PushedFilters: [IsNotNull(profileID), EqualTo(profileID,6)], \| ReadSchema: struct<profileID:int> """.stripMargin ``` 2. When the `Collection`* is empty, and the input is nullable, the logical plan will be simplified to ```scala profileDF.filter( $"profileID".isInCollection(Set())).explain(true) """ \|== Optimized Logical Plan == \|Filter if (isnull(profileID#0)) null else false \|+- Relation[profileID#0] parquet """.stripMargin ``` TODO: 1. For multiple conditions with numbers less than certain thresholds, we should still allow predicate pushdown. 2. Optimize the `In` using `tableswitch` or `lookupswitch` when the numbers of the categories are low, and they are `Int`, `Long`. 3. The default immutable hash trees set is slow for query, and we should do benchmark for using different set implementation for faster query. 4. `filter(if (condition) null else false)` can be optimized to false. ## How was this patch tested? Couple new tests are added. Author: DB Tsai <d_tsai@apple.com> Closes #21797 from dbtsai/optimize-in.		2018-07-17 17:33:52 -07:00
..
catalyst	[SPARK-24402][SQL] Optimize `In` expression when only one element in the collection or collection is empty	2018-07-17 17:33:52 -07:00
core	[SPARK-21590][SS] Window start time should support negative values	2018-07-17 11:25:23 -05:00
hive	[SPARK-24681][SQL] Verify nested column names in Hive metastore	2018-07-17 14:15:30 -07:00
hive-thriftserver	[SPARK-24553][WEB-UI] http 302 fixes for href redirect	2018-06-27 15:36:59 -07:00
create-docs.sh	[MINOR][DOCS] Minor doc fixes related with doc build and uses script dir in SQL doc gen script	2017-08-26 13:56:24 +09:00
gen-sql-markdown.py	[SPARK-21485][FOLLOWUP][SQL][DOCS] Describes examples and arguments separately, and note/since in SQL built-in function documentation	2017-08-05 10:10:56 -07:00
mkdocs.yml	[SPARK-21485][SQL][DOCS] Spark SQL documentation generation for built-in functions	2017-07-26 09:38:51 -07:00
README.md	[MINOR][DOC] Fix some typos and grammar issues	2018-04-06 13:37:08 +08:00

README.md

Spark SQL

This module provides support for executing relational queries expressed in either SQL or the DataFrame/Dataset API.

Spark SQL is broken up into four subprojects:

Catalyst (sql/catalyst) - An implementation-agnostic framework for manipulating trees of relational operators and expressions.
Execution (sql/core) - A query planner / execution engine for translating Catalyst's logical query plans into Spark RDDs. This component also includes a new public interface, SQLContext, that allows users to execute SQL or LINQ statements against existing RDDs and Parquet files.
Hive Support (sql/hive) - Includes an extension of SQLContext called HiveContext that allows users to write queries using a subset of HiveQL and access data from a Hive Metastore using Hive SerDes. There are also wrappers that allow users to run queries that include Hive UDFs, UDAFs, and UDTFs.
HiveServer and CLI support (sql/hive-thriftserver) - Includes support for the SQL CLI (bin/spark-sql) and a HiveServer2 (for JDBC/ODBC) compatible server.

Running sql/create-docs.sh generates SQL documentation for built-in functions under sql/site.