Apache Spark - A unified analytics engine for large-scale data processing

Go to file

Jungtaek Lim 07011eb779 [SPARK-35861][SS] Introduce "prefix match scan" feature on state store ### What changes were proposed in this pull request? This PR proposes to introduce a new feature "prefix match scan" on state store, which enables users of state store (mostly stateful operators) to group the keys into logical groups, and scan the keys in the same group efficiently. For example, if the schema of the key of state store is `[ sessionId \| session.start ]`, we can scan with prefix key which schema is `[ sessionId ]` (leftmost 1 column) and retrieve all key-value pairs in state store which keys are matched with given prefix key. This PR will bring the API changes, though the changes are done in the developer API. * Registering the prefix key We propose to make an explicit change to the init() method of StateStoreProvider, as below: ``` def init( stateStoreId: StateStoreId, keySchema: StructType, valueSchema: StructType, numColsPrefixKey: Int, storeConfs: StateStoreConf, hadoopConf: Configuration): Unit ``` Please note that we remove an unused parameter “keyIndexOrdinal” as well. The parameter is coupled with getRange() which we will remove as well. See below for rationalization. Here we provide the number of columns we take to project the prefix key from the full key. If the operator doesn’t leverage prefix match scan, the value can (and should) be 0, because the state store provider may optimize the underlying storage format which may bring extra overhead. We would like to apply some restrictions on prefix key to simplify the functionality: * Prefix key is a part of the full key. It can’t be the same as the full key. * That said, the full key will be the (prefix key + remaining parts), and both prefix key and remaining parts should have at least one column. * We always take the columns from the leftmost sequentially, like “seq.take(nums)”. * We don’t allow reordering of the columns. * We only guarantee “equality” comparison against prefix keys, and don’t support the prefix “range” scan. * We only support scanning on the keys which match with the prefix key. * E.g. We don’t support the range scan from user A to user B due to technical complexity. That’s the reason we can’t leverage the existing getRange API. As we mentioned, we want to make an explicit change to the init() method of StateStoreProvider which would break backward compatibility, assuming that 3rd party state store providers need to update their code in any way to support prefix match scan. Given RocksDB state store provider is being donated to the OSS and plan to be available in Spark 3.2, the majority of the users would migrate to the built-in state store providers, which would remedy the concerns. * Scanning key-value pairs matched to the prefix key We propose to add a new method to the ReadStateStore (and StateStore by inheritance), as below: ``` def prefixScan(prefixKey: UnsafeRow): Iterator[UnsafeRowPair] ``` We require callers to pass the `prefixKey` which would have the same schema with the registered prefix key schema. In other words, the schema of the parameter `prefixKey` should match to the projection of the prefix key on the full key based on the number of columns for the prefix key. The method contract is clear - the method will return the iterator which will give the key-value pairs whose prefix key is matched with the given prefix key. Callers should only rely on the contract and should not expect any other characteristics based on specific details on the state store provider. In the caller’s point of view, the prefix key is only used for retrieving key-value pairs via prefix match scan. Callers should keep using the full key to do CRUD. Note that this PR also proposes to make a breaking change, removal of getRange(), which is never be implemented properly and hence never be called properly. ### Why are the changes needed? * Introducing prefix match scan feature Currently, the API in state store is only based on key-value data structure. This lacks on advanced data structures like list-like one, which required us to implement the data structure on our own whenever we need it. We had one in stream-stream join, and we were about to have another one in native session window. The custom implementation of data structure based on the state store API tends to be complicated and has to deal with multiple state stores. We decided to enhance the state store API a bit to remove the requirement for native session window to implement its own. From the operator of native session window, it will just need to do prefix scan on group key to retrieve all sessions belonging to the group key. Thanks to adding the feature to the part of state store API, this would enable state store providers to optimize the implementation based on the characteristic. (e.g. We will implement this in RocksDB state store provider via leveraging the characteristic that RocksDB sorts the key by natural order of binary format.) * Removal of getRange API Before introducing this we sought the way to leverage getRange, but it's quite hard to implement efficiently, with respecting its method contract. Spark always calls the method with (None, None) parameter and all the state store providers (including built-in) implement it as just calling iterator(), which is not respecting the method contract. That said, we can replace all getRange() usages to iterator(), and remove the API to remove any confusions/concerns. ### Does this PR introduce _any_ user-facing change? Yes for the end users & maintainers of 3rd party state store provider. They will need to upgrade their state store provider implementations to adopt this change. ### How was this patch tested? Added UT, and also existing UTs to make sure it doesn't break anything. Closes #33038 from HeartSaVioR/SPARK-35861. Authored-by: Jungtaek Lim <kabhwan.opensource@gmail.com> Signed-off-by: Liang-Chi Hsieh <viirya@gmail.com> (cherry picked from commit `094300fa60`) Signed-off-by: Liang-Chi Hsieh <viirya@gmail.com>		2021-07-12 09:07:07 -07:00
.github	[SPARK-35958][CORE] Refactor SparkError.scala to SparkThrowable.java	2021-07-08 23:55:11 +08:00
.idea	[SPARK-35223] Add IssueNavigationLink	2021-04-26 21:51:21 +08:00
assembly	[SPARK-33212][FOLLOWUP] Add hadoop-yarn-server-web-proxy for Hadoop 3.x profile	2021-02-28 16:37:49 -08:00
bin	[SPARK-34688][PYTHON] Upgrade to Py4J 0.10.9.2	2021-03-11 09:51:41 -06:00
binder	[SPARK-35588][PYTHON][DOCS] Merge Binder integration and quickstart notebook for pandas API on Spark	2021-06-24 10:17:22 +09:00
build	[SPARK-35825][INFRA][FOLLOWUP] Increase it in build/mvn script	2021-07-01 22:24:48 -07:00
common	[SPARK-35258][SHUFFLE][YARN] Add new metrics to ExternalShuffleService for better monitoring	2021-06-28 02:36:17 -05:00
conf	[SPARK-35143][SQL][SHELL] Add default log level config for spark-sql	2021-04-23 14:26:19 +09:00
core	[SPARK-36062][PYTHON] Try to capture faulthanlder when a Python worker crashes	2021-07-09 11:31:00 +09:00
data	[SPARK-22666][ML][SQL] Spark datasource for image format	2018-09-05 11:59:00 -07:00
dev	[SPARK-36002][PYTHON] Consolidate tests for data-type-based operations of decimal Series	2021-07-09 14:08:23 +09:00
docs	[SPARK-36089][SQL][DOCS] Update the SQL migration guide about encoding auto-detection of CSV files	2021-07-12 18:54:46 +09:00
examples	[SPARK-35380][SQL] Loading SparkSessionExtensions from ServiceLoader	2021-05-13 16:34:13 +08:00
external	[SPARK-34302][SQL][FOLLOWUP] More code cleanup	2021-07-06 03:43:54 +08:00
graphx	[SPARK-35928][BUILD] Upgrade ASM to 9.1	2021-06-29 10:27:51 -07:00
hadoop-cloud	Revert "[SPARK-36068][BUILD][TEST] No tests in hadoop-cloud run unless hadoop-3.2 profile is activated explicitly"	2021-07-09 18:02:19 +09:00
launcher	[SPARK-33717][LAUNCHER] deprecate spark.launcher.childConectionTimeout	2021-03-26 15:53:52 -05:00
licenses	[SPARK-32435][PYTHON] Remove heapq3 port from Python 3	2020-07-27 20:10:13 +09:00
licenses-binary	[SPARK-35150][ML] Accelerate fallback BLAS with dev.ludovic.netlib	2021-04-27 14:00:59 -05:00
mllib	[SPARK-35678][ML][FOLLOWUP] Revert changes in ANN	2021-06-24 14:02:28 +09:00
mllib-local	[SPARK-35678][ML][FOLLOWUP] softmax support offset and step	2021-06-23 21:03:18 -05:00
project	[SPARK-33996][BUILD][FOLLOW-UP] Match SBT's plugin checkstyle version to Maven's	2021-07-05 18:55:53 +09:00
python	[SPARK-36003][PYTHON] Implement unary operator `invert` of integral ps.Series/Index	2021-07-12 15:10:37 +09:00
R	[SPARK-35968][SQL] Make sure partitions are not too small in AQE partition coalescing	2021-07-02 16:07:46 +08:00
repl	[SPARK-35928][BUILD] Upgrade ASM to 9.1	2021-06-29 10:27:51 -07:00
resource-managers	[SPARK-36067][BUILD][TEST][YARN] YarnClusterSuite fails due to NoClassDefFoundError unless hadoop-3.2 profile is activated explicitly	2021-07-09 15:19:03 +09:00
sbin	[SPARK-34688][PYTHON] Upgrade to Py4J 0.10.9.2	2021-03-11 09:51:41 -06:00
sql	[SPARK-35861][SS] Introduce "prefix match scan" feature on state store	2021-07-12 09:07:07 -07:00
streaming	[SPARK-34520][CORE] Remove unused SecurityManager references	2021-02-24 20:38:03 -08:00
tools	[SPARK-33662][BUILD] Setting version to 3.2.0-SNAPSHOT	2020-12-04 14:10:42 -08:00
.asf.yaml	[MINOR][INFRA] Update a broken link in .asf.yml	2021-01-16 13:42:27 -08:00
.gitattributes	[SPARK-30653][INFRA][SQL] EOL character enforcement for java/scala/xml/py/R files	2020-01-27 10:20:51 -08:00
.gitignore	[SPARK-35842][INFRA] Ignore all .idea folders	2021-06-21 22:07:02 +08:00
appveyor.yml	[SPARK-33757][INFRA][R][FOLLOWUP] Provide more simple solution	2020-12-13 17:27:39 -08:00
CONTRIBUTING.md	[MINOR][DOCS] Tighten up some key links to the project and download pages to use HTTPS	2019-05-21 10:56:42 -07:00
LICENSE	[SPARK-32435][PYTHON] Remove heapq3 port from Python 3	2020-07-27 20:10:13 +09:00
LICENSE-binary	[SPARK-35295][ML] Replace fully com.github.fommil.netlib by dev.ludovic.netlib:2.0	2021-05-12 08:59:36 -05:00
NOTICE	[SPARK-29674][CORE] Update dropwizard metrics to 4.1.x for JDK 9+	2019-11-03 15:13:06 -08:00
NOTICE-binary	[SPARK-29674][CORE] Update dropwizard metrics to 4.1.x for JDK 9+	2019-11-03 15:13:06 -08:00
pom.xml	[SPARK-35992][BUILD] Upgrade ORC to 1.6.9	2021-07-02 09:50:00 -07:00
README.md	[MINOR] Add GitHub Action build status badge to the README	2021-06-17 15:25:24 -07:00
scalastyle-config.xml	[SPARK-35894][BUILD] Introduce new style enforce to not import scala.collection.Seq/IndexedSeq	2021-06-26 09:41:16 +09:00

README.md

Apache Spark

Spark is a unified analytics engine for large-scale data processing. It provides high-level APIs in Scala, Java, Python, and R, and an optimized engine that supports general computation graphs for data analysis. It also supports a rich set of higher-level tools including Spark SQL for SQL and DataFrames, MLlib for machine learning, GraphX for graph processing, and Structured Streaming for stream processing.

https://spark.apache.org/

Online Documentation

You can find the latest Spark documentation, including a programming guide, on the project web page. This README file only contains basic setup instructions.

Building Spark

Spark is built using Apache Maven. To build Spark and its example programs, run:

./build/mvn -DskipTests clean package

(You do not need to do this if you downloaded a pre-built package.)

More detailed documentation is available from the project site, at "Building Spark".

For general development tips, including info on developing Spark using an IDE, see "Useful Developer Tools".

Interactive Scala Shell

The easiest way to start using Spark is through the Scala shell:

./bin/spark-shell

Try the following command, which should return 1,000,000,000:

scala> spark.range(1000 * 1000 * 1000).count()

Interactive Python Shell

Alternatively, if you prefer Python, you can use the Python shell:

./bin/pyspark

And run the following command, which should also return 1,000,000,000:

>>> spark.range(1000 * 1000 * 1000).count()

Example Programs

Spark also comes with several sample programs in the examples directory. To run one of them, use ./bin/run-example <class> [params]. For example:

./bin/run-example SparkPi

will run the Pi example locally.

You can set the MASTER environment variable when running examples to submit examples to a cluster. This can be a mesos:// or spark:// URL, "yarn" to run on YARN, and "local" to run locally with one thread, or "local[N]" to run locally with N threads. You can also use an abbreviated class name if the class is in the examples package. For instance:

MASTER=spark://host:7077 ./bin/run-example SparkPi

Many of the example programs print usage help if no params are given.

Running Tests

Testing first requires building Spark. Once Spark is built, tests can be run using:

./dev/run-tests

Please see the guidance on how to run tests for a module, or individual tests.

There is also a Kubernetes integration test, see resource-managers/kubernetes/integration-tests/README.md

A Note About Hadoop Versions

Spark uses the Hadoop core library to talk to HDFS and other Hadoop-supported storage systems. Because the protocols have changed in different versions of Hadoop, you must build Spark against the same version that your cluster runs.

Please refer to the build documentation at "Specifying the Hadoop Version and Enabling YARN" for detailed guidance on building for a particular distribution of Hadoop, including building for particular Hive and Hive Thriftserver distributions.

Configuration

Please refer to the Configuration Guide in the online documentation for an overview on how to configure Spark.

Contributing

Please review the Contribution to Spark guide for information on how to get started contributing to the project.