ODIn/spark-instrumented-optimizer

Author	SHA1	Message	Date
Nicholas Chammas	0ce5519f17	[SPARK-31153][BUILD] Cleanup several failures in lint-python ### What changes were proposed in this pull request? This PR cleans up several failures -- most of them silent -- in `dev/lint-python`. I don't understand how we haven't been bitten by these yet. Perhaps we've been lucky? Fixes include: * Fix how we compare versions. All the version checks currently in `master` silently fail with: ``` File "<string>", line 2 print(LooseVersion("""2.3.1""") >= LooseVersion("""2.4.0""")) ^ IndentationError: unexpected indent ``` Another problem is that `distutils.version` is undocumented and unsupported. * Fix some basic bugs. e.g. We have an incorrect reference to `$PYDOCSTYLEBUILD`, which doesn't exist, which was causing the doc style test to silently fail with: ``` ./dev/lint-python: line 193: --version: command not found ``` * Stop suppressing error output! It's hiding problems and serves no purpose here. ### Why are the changes needed? `lint-python` is part of our CI build and is currently doing any combination of the following: silently failing; incorrectly skipping tests; incorrectly downloading libraries when a suitable library is already available. ### Does this PR introduce any user-facing change? No. ### How was this patch tested? Lots of manual testing with `set -x` enabled. Closes #27910 from nchammas/SPARK-31153-lint-python. Authored-by: Nicholas Chammas <nicholas.chammas@liveramp.com> Signed-off-by: HyukjinKwon <gurwls223@apache.org>	2020-03-15 13:09:35 +09:00
gatorsmile	4d4c00c1b5	[SPARK-31151][SQL][DOC] Reorganize the migration guide of SQL ### What changes were proposed in this pull request? The current migration guide of SQL is too long for most readers to find the needed info. This PR is to group the items in the migration guide of Spark SQL based on the corresponding components. Note. This PR does not change the contents of the migration guides. Attached figure is the screenshot after the change. ![screencapture-127-0-0-1-4000-sql-migration-guide-html-2020-03-14-12_00_40](https://user-images.githubusercontent.com/11567269/76688626-d3010200-65eb-11ea-9ce7-265bc90ebb2c.png) ### Why are the changes needed? The current migration guide of SQL is too long for most readers to find the needed info. ### Does this PR introduce any user-facing change? No ### How was this patch tested? N/A Closes #27909 from gatorsmile/migrationGuideReorg. Authored-by: gatorsmile <gatorsmile@gmail.com> Signed-off-by: Takeshi Yamamuro <yamamuro@apache.org>	2020-03-15 07:35:20 +09:00
HyukjinKwon	9628aca68b	[MINOR][DOCS] Fix [[...]] to `...` and <code>...</code> in documentation ### What changes were proposed in this pull request? Before: - ![Screen Shot 2020-03-13 at 1 19 12 PM](https://user-images.githubusercontent.com/6477701/76589452-7c34f300-652d-11ea-9da7-3754f8575796.png) - ![Screen Shot 2020-03-13 at 1 19 24 PM](https://user-images.githubusercontent.com/6477701/76589455-7d662000-652d-11ea-9dbe-f5fe10d1e7ad.png) - ![Screen Shot 2020-03-13 at 1 19 03 PM](https://user-images.githubusercontent.com/6477701/76589449-7b03c600-652d-11ea-8e99-dbe47f561f9c.png) After: - ![Screen Shot 2020-03-13 at 1 17 37 PM](https://user-images.githubusercontent.com/6477701/76589437-74754e80-652d-11ea-99f5-14fb4761f915.png) - ![Screen Shot 2020-03-13 at 1 17 46 PM](https://user-images.githubusercontent.com/6477701/76589442-76d7a880-652d-11ea-8c10-53e595421081.png) - ![Screen Shot 2020-03-13 at 1 18 15 PM](https://user-images.githubusercontent.com/6477701/76589443-7808d580-652d-11ea-9b1b-e5d11d638335.png) ### Why are the changes needed? To render the code block properly in the documentation ### Does this PR introduce any user-facing change? Yes, code rendering in documentation. ### How was this patch tested? Manually built the doc via `SKIP_API=1 jekyll build`. Closes #27899 from HyukjinKwon/minor-docss. Authored-by: HyukjinKwon <gurwls223@apache.org> Signed-off-by: Dongjoon Hyun <dongjoon@apache.org>	2020-03-13 16:44:23 -07:00
Shixiong Zhu	1ddf44dfca	[SPARK-31144][SQL] Wrap Error with QueryExecutionException to notify QueryExecutionListener ### What changes were proposed in this pull request? This PR manually reverts changes in #25292 and then wraps java.lang.Error with `QueryExecutionException` to notify `QueryExecutionListener` to send it to `QueryExecutionListener.onFailure` which only accepts `Exception`. The bug fix PR for 2.4 is #27904. It needs a separate PR because the touched codes were changed a lot. ### Why are the changes needed? Avoid API changes and fix a bug. ### Does this PR introduce any user-facing change? Yes. Reverting an API change happening in 3.0. QueryExecutionListener APIs will be the same as 2.4. ### How was this patch tested? The new added test. Closes #27907 from zsxwing/SPARK-31144. Authored-by: Shixiong Zhu <zsxwing@gmail.com> Signed-off-by: Dongjoon Hyun <dongjoon@apache.org>	2020-03-13 15:55:29 -07:00
Dale Clarke	2a4fed0443	[SPARK-30654][WEBUI] Bootstrap4 WebUI upgrade ### What changes were proposed in this pull request? Spark's Web UI is using an older version of Bootstrap (v. 2.3.2) for the portal pages. Bootstrap 2.x was moved to EOL in Aug 2013 and Bootstrap 3.x was moved to EOL in July 2019 (https://github.com/twbs/release). Older versions of Bootstrap are also getting flagged in security scans for various CVEs: https://snyk.io/vuln/SNYK-JS-BOOTSTRAP-72889 https://snyk.io/vuln/SNYK-JS-BOOTSTRAP-173700 https://snyk.io/vuln/npm:bootstrap:20180529 https://snyk.io/vuln/npm:bootstrap:20160627 I haven't validated each CVE, but it would be nice to resolve any potential issues and get on a supported release. The bad news is that there have been quite a few changes between Bootstrap 2 and Bootstrap 4. I've tried updating the library, refactoring/tweaking the CSS and JS to maintain a similar appearance and functionality, and testing the UI for functionality and appearance. This is a fairly large change so I'm sure additional testing and fixes will be needed. ### How was this patch tested? This has been manually tested, but there is a ton of functionality and there are many pages and detail pages so it is very possible bugs introduced from the upgrade were missed. Additional testing and feedback is welcomed. If it appears a whole page was missed let me know and I'll take a pass at addressing that page/section. Closes #27370 from clarkead/bootstrap4-core-upgrade. Authored-by: Dale Clarke <a.dale.clarke@gmail.com> Signed-off-by: Gengliang Wang <gengliang.wang@databricks.com>	2020-03-13 15:24:48 -07:00
Kousuke Saruta	680981587d	[SPARK-31004][WEBUI][SS] Show message for empty Streaming Queries instead of empty timelines and histograms ### What changes were proposed in this pull request? `StreamingQueryStatisticsPage` shows a message "No visualization information available because there is no batches" instead of showing empty timelines and histograms for empty streaming queries. [Before this change applied] ![before-fix-for-empty-streaming-query](https://user-images.githubusercontent.com/4736016/75642391-b32e1d80-5c7e-11ea-9c07-e2f0f1b5b4f9.png) [After this change applied] ![after-fix-for-empty-streaming-query2](https://user-images.githubusercontent.com/4736016/75694583-1904be80-5cec-11ea-9b13-dc7078775188.png) ### Why are the changes needed? Empty charts are ugly and a little bit confusing. It's better to clearly say "No visualization information available". Also, this change fixes a JS error shown in the capture above. This error occurs because `drawTimeline` in `streaming-page.js` is called even though `formattedDate` will be `undefined` for empty streaming queries. ### Does this PR introduce any user-facing change? Yes. screen captures are shown above. ### How was this patch tested? Manually tested by creating an empty streaming query like as follows. ``` val df = spark.readStream.format("socket").options(Map("host"->"<non-existing hostname>", "port"->"...")).load df.writeStream.format("console").start ``` This streaming query will fail because of `non-existing hostname` and has no batches. Closes #27755 from sarutak/fix-for-empty-batches. Authored-by: Kousuke Saruta <sarutak@oss.nttdata.com> Signed-off-by: Gengliang Wang <gengliang.wang@databricks.com>	2020-03-13 12:58:49 -07:00
Gengliang Wang	0f463258c2	[SPARK-31128][WEBUI] Fix Uncaught TypeError in streaming statistics page ### What changes were proposed in this pull request? There is a minor issue in https://github.com/apache/spark/pull/26201 In the streaming statistics page, there is such error ``` streaming-page.js:211 Uncaught TypeError: Cannot read property 'top' of undefined at SVGCircleElement.<anonymous> (streaming-page.js:211) at SVGCircleElement.__onclick (d3.min.js:1) ``` in the console after clicking the timeline graph. ![image](https://user-images.githubusercontent.com/1097932/76479745-14b26280-63ca-11ea-9079-0065321795f9.png) This PR is to fix it. ### Why are the changes needed? Fix the error of javascript execution. ### Does this PR introduce any user-facing change? No, the error shows up in the console. ### How was this patch tested? Manual test. Closes #27883 from gengliangwang/fixSelector. Authored-by: Gengliang Wang <gengliang.wang@databricks.com> Signed-off-by: Gengliang Wang <gengliang.wang@databricks.com>	2020-03-12 20:01:17 -07:00
gatorsmile	1c8526dc87	[SPARK-28093][FOLLOW-UP] Remove migration guide of TRIM changes ### What changes were proposed in this pull request? Since we reverted the original change in https://github.com/apache/spark/pull/27540, this PR is to remove the corresponding migration guide made in the commit https://github.com/apache/spark/pull/24948 ### Why are the changes needed? N/A ### Does this PR introduce any user-facing change? N/A ### How was this patch tested? N/A Closes #27896 from gatorsmile/SPARK-28093Followup. Authored-by: gatorsmile <gatorsmile@gmail.com> Signed-off-by: HyukjinKwon <gurwls223@apache.org>	2020-03-13 11:45:59 +09:00
Gabor Somogyi	231e65092f	[SPARK-30874][SQL] Support Postgres Kerberos login in JDBC connector ### What changes were proposed in this pull request? When loading DataFrames from JDBC datasource with Kerberos authentication, remote executors (yarn-client/cluster etc. modes) fail to establish a connection due to lack of Kerberos ticket or ability to generate it. This is a real issue when trying to ingest data from kerberized data sources (SQL Server, Oracle) in enterprise environment where exposing simple authentication access is not an option due to IT policy issues. In this PR I've added Postgres support (other supported databases will come in later PRs). What this PR contains: * Added `keytab` and `principal` JDBC options * Added `ConnectionProvider` trait and it's impementations: * `BasicConnectionProvider` => unsecure connection * `PostgresConnectionProvider` => postgres secure connection * Added `ConnectionProvider` tests * Added `PostgresKrbIntegrationSuite` docker integration test * Created `SecurityUtils` to concentrate re-usable security related functionalities * Documentation ### Why are the changes needed? Missing JDBC kerberos support. ### Does this PR introduce any user-facing change? Yes, 2 additional JDBC options added: * keytab * principal If both provided then Spark does kerberos authentication. ### How was this patch tested? To demonstrate the functionality with a standalone application I've created this repository: https://github.com/gaborgsomogyi/docker-kerberos * Additional + existing unit tests * Additional docker integration test * Test on cluster manually * `SKIP_API=1 jekyll build` Closes #27637 from gaborgsomogyi/SPARK-30874. Authored-by: Gabor Somogyi <gabor.g.somogyi@gmail.com> Signed-off-by: Marcelo Vanzin <vanzin@apache.org>	2020-03-12 19:04:35 -07:00
Wenchen Fan	b27b3c91f1	[SPARK-31090][SPARK-25457] Revert "IntegralDivide returns data type of the operands" ### What changes were proposed in this pull request? This reverts commit `47d6e80a2e`. ### Why are the changes needed? There is no standard requiring that `div` must return the type of the operand, and always returning long type looks fine. This is kind of a cosmetic change and we should avoid it if it breaks existing queries. This is similar to reverting TRIM function parameter order change. ### Does this PR introduce any user-facing change? Yes, change the behavior of `div` back to be the same as 2.4. ### How was this patch tested? N/A Closes #27835 from cloud-fan/revert2. Authored-by: Wenchen Fan <wenchen@databricks.com> Signed-off-by: HyukjinKwon <gurwls223@apache.org>	2020-03-13 10:47:36 +09:00
Kent Yao	fbc9dc7e9d	[SPARK-31129][SQL][TESTS] Fix IntervalBenchmark and DateTimeBenchmark ### What changes were proposed in this pull request? This PR aims to recover `IntervalBenchmark` and `DataTimeBenchmark` due to banning intervals as output. ### Why are the changes needed? This PR recovers the benchmark suite. ### Does this PR introduce any user-facing change? No. ### How was this patch tested? Manually, re-run the benchmark. Closes #27885 from yaooqinn/SPARK-31111-2. Authored-by: Kent Yao <yaooqinn@hotmail.com> Signed-off-by: Dongjoon Hyun <dongjoon@apache.org>	2020-03-12 12:59:29 -07:00
Kent Yao	7b4b29e8d9	[SPARK-31131][SQL] Remove the unnecessary config spark.sql.legacy.timeParser.enabled ### What changes were proposed in this pull request? spark.sql.legacy.timeParser.enabled should be removed from SQLConf and the migration guide spark.sql.legacy.timeParsePolicy is the right one ### Why are the changes needed? fix doc ### Does this PR introduce any user-facing change? no ### How was this patch tested? Pass the jenkins Closes #27889 from yaooqinn/SPARK-31131. Authored-by: Kent Yao <yaooqinn@hotmail.com> Signed-off-by: Dongjoon Hyun <dongjoon@apache.org>	2020-03-12 09:24:49 -07:00
Dongjoon Hyun	972e23d181	[SPARK-31130][BUILD] Use the same version of `commons-io` in SBT ### What changes were proposed in this pull request? This PR (SPARK-31130) aims to pin `Commons IO` version to `2.4` in SBT build like Maven build. ### Why are the changes needed? [HADOOP-15261](https://issues.apache.org/jira/browse/HADOOP-15261) upgraded `commons-io` from 2.4 to 2.5 at Apache Hadoop 3.1. In `Maven`, Apache Spark always uses `Commons IO 2.4` based on `pom.xml`. ``` $ git grep commons-io.version pom.xml: <commons-io.version>2.4</commons-io.version> pom.xml: <version>${commons-io.version}</version> ``` However, `SBT` choose `2.5`. branch-3.0 ``` $ build/sbt -Phadoop-3.2 "core/dependencyTree" \| grep commons-io:commons-io \| head -n1 [info] \| \| +-commons-io:commons-io:2.5 ``` branch-2.4 ``` $ build/sbt -Phadoop-3.1 "core/dependencyTree" \| grep commons-io:commons-io \| head -n1 [info] \| \| +-commons-io:commons-io:2.5 ``` ### Does this PR introduce any user-facing change? No. ### How was this patch tested? Pass the Jenkins with `[test-hadoop3.2]` (the default PR Builder is `SBT`) and manually do the following locally. ``` build/sbt -Phadoop-3.2 "core/dependencyTree" \| grep commons-io:commons-io \| head -n1 ``` Closes #27886 from dongjoon-hyun/SPARK-31130. Authored-by: Dongjoon Hyun <dongjoon@apache.org> Signed-off-by: Dongjoon Hyun <dongjoon@apache.org>	2020-03-12 09:06:29 -07:00
Wenchen Fan	77c49cb702	[SPARK-31124][SQL] change the default value of minPartitionNum in AQE ### What changes were proposed in this pull request? AQE has a perf regression when using the default settings: if we coalesce the shuffle partitions into one or few partitions, we may leave many CPU cores idle and the perf is worse than with AQE off (which leverages all CPU cores). Technically, this is not a bad thing. If there are many queries running at the same time, it's better to coalesce shuffle partitions into fewer partitions. However, the default settings of AQE should try to avoid any perf regression as possible as we can. This PR changes the default value of minPartitionNum when coalescing shuffle partitions, to be `SparkContext.defaultParallelism`, so that AQE can leverage all the CPU cores. ### Why are the changes needed? avoid AQE perf regression ### Does this PR introduce any user-facing change? No ### How was this patch tested? existing tests Closes #27879 from cloud-fan/aqe. Authored-by: Wenchen Fan <wenchen@databricks.com> Signed-off-by: Wenchen Fan <wenchen@databricks.com>	2020-03-12 21:28:24 +08:00
yi.wu	feb9b9e771	[SPARK-31010][SQL][FOLLOW-UP] Give an example for typed Scala UDF in error message ### What changes were proposed in this pull request? In the error message, adding an example for typed Scala UDF. ### Why are the changes needed? Help user to know how to migrate to typed Scala UDF. ### Does this PR introduce any user-facing change? No, it's a new error message in Spark 3.0. ### How was this patch tested? Pass Jenkins. Closes #27884 from Ngone51/spark_31010_followup. Authored-by: yi.wu <yi.wu@databricks.com> Signed-off-by: HyukjinKwon <gurwls223@apache.org>	2020-03-12 21:16:02 +09:00
Kent Yao	18f2730874	[SPARK-31066][SQL][TEST-HIVE1.2] Disable useless and uncleaned hive SessionState initialization parts ### What changes were proposed in this pull request? As a common usage and according to the spark doc, users may often just copy their `hive-site.xml` to Spark directly from hive projects. Sometimes, the config file is not that clean for spark and may cause some side effects. for example, `hive.session.history.enabled` will create a log for the hive jobs but useless for spark and also it will not be deleted on JVM exit. this pr 1) disable `hive.session.history.enabled` explicitly to disable creating `hive_job_log` file, e.g. ``` Hive history file=/var/folders/01/h81cs4sn3dq2dd_k4j6fhrmc0000gn/T//kentyao/hive_job_log_79c63b29-95a4-4935-a9eb-2d89844dfe4f_493861201.txt ``` 2) set `hive.execution.engine` to `spark` explicitly in case the config is `tez` and casue uneccesary problem like this: ``` Exception in thread "main" java.lang.NoClassDefFoundError: org/apache/tez/dag/api/SessionNotRunning at org.apache.hadoop.hive.ql.session.SessionState.start(SessionState.java:529) ``` ### Why are the changes needed? reduce overhead of internal complexity and users' hive cognitive load for running spark ### Does this PR introduce any user-facing change? yes, `hive_job_log` file will not be created even enabled, and will not try to initialize tez kinds of stuff ### How was this patch tested? add ut and verify manually Closes #27827 from yaooqinn/SPARK-31066. Authored-by: Kent Yao <yaooqinn@hotmail.com> Signed-off-by: Wenchen Fan <wenchen@databricks.com>	2020-03-12 18:13:52 +08:00
Jungtaek Lim (HeartSaVioR)	3946b24328	[SPARK-31011][CORE] Log better message if SIGPWR is not supported while setting up decommission ### What changes were proposed in this pull request? This patch changes to log better message (at least relevant to decommission) when registering signal handler for SIGPWR fails. SIGPWR is non-POSIX and not all unix-like OS support it; we can easily find the case, macOS. ### Why are the changes needed? Spark already logs message on failing to register handler for SIGPWR, but the error message is too general which doesn't give the information of the impact. End users should be noticed that failing to register handler for SIGPWR effectively "disables" the feature of decommission. ### Does this PR introduce any user-facing change? No. ### How was this patch tested? Manually tested via running standalone master/worker in macOS 10.14.6, with `spark.worker.decommission.enabled= true`, and submit an example application to run executors. (NOTE: the message may be different a bit, as the message can be updated in review phase.) For worker log: ``` 20/03/06 17:19:13 INFO Worker: Registering SIGPWR handler to trigger decommissioning. 20/03/06 17:19:13 INFO SignalUtils: Registering signal handler for PWR 20/03/06 17:19:13 WARN SignalUtils: Failed to register SIGPWR - disabling worker decommission. java.lang.IllegalArgumentException: Unknown signal: PWR at java.base/jdk.internal.misc.Signal.<init>(Signal.java:148) at jdk.unsupported/sun.misc.Signal.<init>(Signal.java:139) at org.apache.spark.util.SignalUtils$.$anonfun$registerSignal$1(SignalUtils.scala:95) at scala.collection.mutable.HashMap.getOrElseUpdate(HashMap.scala:86) at org.apache.spark.util.SignalUtils$.registerSignal(SignalUtils.scala:93) at org.apache.spark.util.SignalUtils$.register(SignalUtils.scala:81) at org.apache.spark.deploy.worker.Worker.<init>(Worker.scala:73) at org.apache.spark.deploy.worker.Worker$.startRpcEnvAndEndpoint(Worker.scala:887) at org.apache.spark.deploy.worker.Worker$.main(Worker.scala:855) at org.apache.spark.deploy.worker.Worker.main(Worker.scala) ``` For executor: ``` 20/03/06 17:21:52 INFO CoarseGrainedExecutorBackend: Registering PWR handler. 20/03/06 17:21:52 INFO SignalUtils: Registering signal handler for PWR 20/03/06 17:21:52 WARN SignalUtils: Failed to register SIGPWR - disabling decommission feature. java.lang.IllegalArgumentException: Unknown signal: PWR at java.base/jdk.internal.misc.Signal.<init>(Signal.java:148) at jdk.unsupported/sun.misc.Signal.<init>(Signal.java:139) at org.apache.spark.util.SignalUtils$.$anonfun$registerSignal$1(SignalUtils.scala:95) at scala.collection.mutable.HashMap.getOrElseUpdate(HashMap.scala:86) at org.apache.spark.util.SignalUtils$.registerSignal(SignalUtils.scala:93) at org.apache.spark.util.SignalUtils$.register(SignalUtils.scala:81) at org.apache.spark.executor.CoarseGrainedExecutorBackend.onStart(CoarseGrainedExecutorBackend.scala:86) at org.apache.spark.rpc.netty.Inbox.$anonfun$process$1(Inbox.scala:120) at org.apache.spark.rpc.netty.Inbox.safelyCall(Inbox.scala:203) at org.apache.spark.rpc.netty.Inbox.process(Inbox.scala:100) at org.apache.spark.rpc.netty.MessageLoop.org$apache$spark$rpc$netty$MessageLoop$$receiveLoop(MessageLoop.scala:75) at org.apache.spark.rpc.netty.MessageLoop$$anon$1.run(MessageLoop.scala:41) at java.base/java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:515) at java.base/java.util.concurrent.FutureTask.run(FutureTask.java:264) at java.base/java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1128) at java.base/java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:628) at java.base/java.lang.Thread.run(Thread.java:834) ``` Closes #27832 from HeartSaVioR/SPARK-31011. Authored-by: Jungtaek Lim (HeartSaVioR) <kabhwan.opensource@gmail.com> Signed-off-by: Dongjoon Hyun <dongjoon@apache.org>	2020-03-11 20:27:00 -07:00
zhengruifeng	7f3c8fa42e	[SPARK-31032][ML] GMM compute summary and update distributions in one job ### What changes were proposed in this pull request? 1, compute summary and update distributions in one pass; 2, remove logic related to check `shouldDistributeGaussians` ### Why are the changes needed? In current impl, GMM need to trigger two jobs at one iteration: 1, one to compute summary; 2, if `shouldDistributeGaussians = ((k - 1.0) / k) * numFeatures > 25.0`, trigger another to update distributions; `shouldDistributeGaussians` is almost true in practice, since numFeatures is likely to be greater than 25. We can use only one job to impl above computation, by following the logic in `KMeans`: using `reduceByKey` to compute statistics for each center ### Does this PR introduce any user-facing change? No ### How was this patch tested? existing testsuites Closes #27784 from zhengruifeng/gmm_avoid_distri_gaussian. Authored-by: zhengruifeng <ruifengz@foxmail.com> Signed-off-by: zhengruifeng <ruifengz@foxmail.com>	2020-03-12 11:21:30 +08:00
Dongjoon Hyun	614323d326	[SPARK-31126][SS] Upgrade Kafka to 2.4.1 ### What changes were proposed in this pull request? This PR (SPARK-31126) aims to upgrade Kafka library to bring a client-side bug fix like KAFKA-8933 ### Why are the changes needed? The following is the full release note. - https://downloads.apache.org/kafka/2.4.1/RELEASE_NOTES.html ### Does this PR introduce any user-facing change? No ### How was this patch tested? Pass the Jenkins with the existing test. Closes #27881 from dongjoon-hyun/SPARK-KAFKA-2.4.1. Authored-by: Dongjoon Hyun <dongjoon@apache.org> Signed-off-by: Dongjoon Hyun <dongjoon@apache.org>	2020-03-11 19:26:15 -07:00
beliefer	bd2b3f9132	[SPARK-30911][CORE][DOC] Add version information to the configuration of Status ### What changes were proposed in this pull request? 1.Add version information to the configuration of `Status`. 2.Update the docs of `Status`. 3.By the way supplementary documentation about https://github.com/apache/spark/pull/27847 I sorted out some information show below. Item name \| Since version \| JIRA ID \| Commit ID \| Note -- \| -- \| -- \| -- \| -- spark.appStateStore.asyncTracking.enable \| 2.3.0 \| SPARK-20653 \| 772e4648d95bda3353723337723543c741ea8476#diff-9ab674b7af7b2097f7d28cb6f5fd1e8c \| spark.ui.liveUpdate.period \| 2.3.0 \| SPARK-20644 \| c7f38e5adb88d43ef60662c5d6ff4e7a95bff580#diff-9ab674b7af7b2097f7d28cb6f5fd1e8c \| spark.ui.liveUpdate.minFlushPeriod \| 2.4.2 \| SPARK-27394 \| a8a2ba11ac10051423e58920062b50f328b06421#diff-9ab674b7af7b2097f7d28cb6f5fd1e8c \| spark.ui.retainedJobs \| 1.2.0 \| SPARK-2321 \| 9530316887612dca060a128fca34dd5a6ab2a9a9#diff-1f32bcb61f51133bd0959a4177a066a5 \| spark.ui.retainedStages \| 0.9.0 \| None \| 112c0a1776bbc866a1026a9579c6f72f293414c4#diff-1f32bcb61f51133bd0959a4177a066a5 \| 0.9.0-incubating-SNAPSHOT spark.ui.retainedTasks \| 2.0.1 \| SPARK-15083 \| 55db26245d69bb02b7d7d5f25029b1a1cd571644#diff-6bdad48cfc34314e89599655442ff210 \| spark.ui.retainedDeadExecutors \| 2.0.0 \| SPARK-7729 \| 9f4263392e492b5bc0acecec2712438ff9a257b7#diff-a0ba36f9b1f9829bf3c4689b05ab6cf2 \| spark.ui.dagGraph.retainedRootRDDs \| 2.1.0 \| SPARK-17171 \| cc87280fcd065b01667ca7a59a1a32c7ab757355#diff-3f492c527ea26679d4307041b28455b8 \| spark.metrics.appStatusSource.enabled \| 3.0.0 \| SPARK-30060 \| 60f20e5ea2000ab8f4a593b5e4217fd5637c5e22#diff-9f796ae06b0272c1f0a012652a5b68d0 \| ### Why are the changes needed? Supplemental configuration version information. ### Does this PR introduce any user-facing change? No ### How was this patch tested? Exists UT Closes #27848 from beliefer/add-version-to-status-config. Lead-authored-by: beliefer <beliefer@163.com> Co-authored-by: Jiaan Geng <beliefer@163.com> Signed-off-by: HyukjinKwon <gurwls223@apache.org>	2020-03-12 11:03:47 +09:00
beliefer	1cd80fa9fa	[SPARK-31109][MESOS][DOC] Add version information to the configuration of Mesos ### What changes were proposed in this pull request? Add version information to the configuration of `Mesos`. I sorted out some information show below. Item name \| Since version \| JIRA ID \| Commit ID \| Note -- \| -- \| -- \| -- \| -- spark.mesos.$taskType.secret.names \| 2.3.0 \| SPARK-22131 \| 5415963d2caaf95604211419ffc4e29fff38e1d7#diff-91e6e5f871160782dc50d4060d6faea3 \| spark.mesos.$taskType.secret.values \| 2.3.0 \| SPARK-22131 \| 5415963d2caaf95604211419ffc4e29fff38e1d7#diff-91e6e5f871160782dc50d4060d6faea3 \| spark.mesos.$taskType.secret.envkeys \| 2.3.0 \| SPARK-22131 \| 5415963d2caaf95604211419ffc4e29fff38e1d7#diff-91e6e5f871160782dc50d4060d6faea3 \| spark.mesos.$taskType.secret.filenames \| 2.3.0 \| SPARK-22131 \| 5415963d2caaf95604211419ffc4e29fff38e1d7#diff-91e6e5f871160782dc50d4060d6faea3 \| spark.mesos.principal \| 1.5.0 \| SPARK-6284 \| d86bbb4e286f16f77ba125452b07827684eafeed#diff-02a6d899f7a529eb7cfbb12182a110b0 \| spark.mesos.principal.file \| 2.4.0 \| SPARK-16501 \| 7f10cf83f311526737fc96d5bb8281d12e41932f#diff-daf48dabbe58afaeed8787751750b01d \| spark.mesos.secret \| 1.5.0 \| SPARK-6284 \| d86bbb4e286f16f77ba125452b07827684eafeed#diff-02a6d899f7a529eb7cfbb12182a110b0 \| spark.mesos.secret.file \| 2.4.0 \| SPARK-16501 \| 7f10cf83f311526737fc96d5bb8281d12e41932f#diff-daf48dabbe58afaeed8787751750b01d \| spark.shuffle.cleaner.interval \| 2.0.0 \| SPARK-12583 \| 310981d49a332bd329303f610b150bbe02cf5f87#diff-2fafefee94f2a2023ea9765536870258 \| spark.mesos.dispatcher.webui.url \| 2.0.0 \| SPARK-13492 \| a4a0addccffb7cd0ece7947d55ce2538afa54c97#diff-f541460c7a74cee87cbb460b3b01665e \| spark.mesos.dispatcher.historyServer.url \| 2.1.0 \| SPARK-16809 \| 62e62124419f3fa07b324f5e42feb2c5b4fde715#diff-3779e2035d9a09fa5f6af903925b9512 \| spark.mesos.driver.labels \| 2.3.0 \| SPARK-21000 \| 8da3f7041aafa71d7596b531625edb899970fec2#diff-91e6e5f871160782dc50d4060d6faea3 \| spark.mesos.driver.webui.url \| 2.0.0 \| SPARK-13492 \| a4a0addccffb7cd0ece7947d55ce2538afa54c97#diff-e3a5e67b8de2069ce99801372e214b8e \| spark.mesos.driver.failoverTimeout \| 2.3.0 \| SPARK-21456 \| c42ef953343073a50ef04c5ce848b574ff7f2238#diff-91e6e5f871160782dc50d4060d6faea3 \| spark.mesos.network.name \| 2.1.0 \| SPARK-18232 \| d89bfc92302424406847ac7a9cfca714e6b742fc#diff-ab5bf34f1951a8f7ea83c9456a6c3ab7 \| spark.mesos.network.labels \| 2.3.0 \| SPARK-21694 \| ce0d3bb377766bdf4df7852272557ae846408877#diff-91e6e5f871160782dc50d4060d6faea3 \| spark.mesos.driver.constraints \| 2.2.1 \| SPARK-19606 \| f6ee3d90d5c299e67ae6e2d553c16c0d9759d4b5#diff-91e6e5f871160782dc50d4060d6faea3 \| spark.mesos.driver.frameworkId \| 2.1.0 \| SPARK-16809 \| 62e62124419f3fa07b324f5e42feb2c5b4fde715#diff-02a6d899f7a529eb7cfbb12182a110b0 \| spark.executor.uri \| 0.8.0 \| None \| 46eecd110a4017ea0c86cbb1010d0ccd6a5eb2ef#diff-a885e7df97790e9b59c21c63353e7476 \| spark.mesos.proxy.baseURL \| 2.3.0 \| SPARK-13041 \| 663f30d14a0c9219e07697af1ab56e11a714d9a6#diff-0b9b4e122eb666155aa189a4321a6ca8 \| spark.mesos.coarse \| 0.6.0 \| None \| 63051dd2bcc4bf09d413ff7cf89a37967edc33ba#diff-eaf125f56ce786d64dcef99cf446a751 \| spark.mesos.coarse.shutdownTimeout \| 2.0.0 \| SPARK-12330 \| c756bda477f458ba4aad7fdb2026263507e0ad9b#diff-d425d35aa23c47a62fbb538554f2f2cf \| spark.mesos.maxDrivers \| 1.4.0 \| SPARK-5338 \| 53befacced828bbac53c6e3a4976ec3f036bae9e#diff-b964c449b99c51f0a5fd77270b2951a4 \| spark.mesos.retainedDrivers \| 1.4.0 \| SPARK-5338 \| 53befacced828bbac53c6e3a4976ec3f036bae9e#diff-b964c449b99c51f0a5fd77270b2951a4 \| spark.mesos.cluster.retry.wait.max \| 1.4.0 \| SPARK-5338 \| 53befacced828bbac53c6e3a4976ec3f036bae9e#diff-b964c449b99c51f0a5fd77270b2951a4 \| spark.mesos.fetcherCache.enable \| 2.1.0 \| SPARK-15994 \| e34b4e12673fb76c92f661d7c03527410857a0f8#diff-772ea7311566edb25f11a4c4f882179a \| spark.mesos.appJar.local.resolution.mode \| 2.4.0 \| SPARK-24326 \| 22df953f6bb191858053eafbabaa5b3ebca29f56#diff-6e4d0a0445975f03f975fdc1e3d80e49 \| spark.mesos.rejectOfferDuration \| 2.2.0 \| SPARK-19702 \| 2e30c0b9bcaa6f7757bd85d1f1ec392d5f916f83#diff-daf48dabbe58afaeed8787751750b01d \| spark.mesos.rejectOfferDurationForUnmetConstraints \| 1.6.0 \| SPARK-10471 \| 74f50275e429e649212928a9f36552941b862edc#diff-02a6d899f7a529eb7cfbb12182a110b0 \| spark.mesos.rejectOfferDurationForReachedMaxCores \| 2.0.0 \| SPARK-13001 \| 1e7d9bfb5a41f5c2479ab3b4d4081f00bf00bd31#diff-02a6d899f7a529eb7cfbb12182a110b0 \| spark.mesos.uris \| 1.5.0 \| SPARK-8798 \| a2f805729b401c68b60bd690ad02533b8db57b58#diff-e3a5e67b8de2069ce99801372e214b8e \| spark.mesos.executor.home \| 1.1.1 \| SPARK-3264 \| 069ecfef02c4af69fc0d3755bd78be321b68b01d#diff-e3a5e67b8de2069ce99801372e214b8e \| spark.mesos.mesosExecutor.cores \| 1.4.0 \| SPARK-6350 \| 6fbeb82e13db7117d8f216e6148632490a4bc5be#diff-e3a5e67b8de2069ce99801372e214b8e \| spark.mesos.extra.cores \| 0.6.0 \| None \| 2d761e3353651049f6707c74bb5ffdd6e86f6f35#diff-37af8c6e3634f97410ade813a5172621 \| spark.mesos.executor.memoryOverhead \| 1.1.1 \| SPARK-3535 \| 6f150978477830bbc14ba983786dd2bce12d1fe2#diff-6b498f5407d10e848acac4a1b182457c \| spark.mesos.executor.docker.image \| 1.4.0 \| SPARK-2691 \| 8f50a07d2188ccc5315d979755188b1e5d5b5471#diff-e3a5e67b8de2069ce99801372e214b8e \| spark.mesos.executor.docker.forcePullImage \| 2.1.0 \| SPARK-15271 \| 978cd5f125eb5a410bad2e60bf8385b11cf1b978#diff-0dd025320c7ecda2ea310ed7172d7f5a \| spark.mesos.executor.docker.portmaps \| 1.4.0 \| SPARK-7373 \| 226033cfffa2f37ebaf8bc2c653f094e91ef0c9b#diff-b964c449b99c51f0a5fd77270b2951a4 \| spark.mesos.executor.docker.parameters \| 2.2.0 \| SPARK-19740 \| a888fed3099e84c2cf45e9419f684a3658ada19d#diff-4139e6605a8c7f242f65cde538770c99 \| spark.mesos.executor.docker.volumes \| 1.4.0 \| SPARK-7373 \| 226033cfffa2f37ebaf8bc2c653f094e91ef0c9b#diff-b964c449b99c51f0a5fd77270b2951a4 \| spark.mesos.gpus.max \| 2.1.0 \| SPARK-14082 \| 29f186bfdf929b1e8ffd8e33ee37b76d5dc5af53#diff-d427ee890b913c5a7056be21eb4f39d7 \| spark.mesos.task.labels \| 2.2.0 \| SPARK-20085 \| c8fc1f3badf61bcfc4bd8eeeb61f73078ca068d1#diff-387c5d0c916278495fc28420571adf9e \| spark.mesos.constraints \| 1.5.0 \| SPARK-6707 \| 1165b17d24cdf1dbebb2faca14308dfe5c2a652c#diff-e3a5e67b8de2069ce99801372e214b8e \| spark.mesos.containerizer \| 2.1.0 \| SPARK-16637 \| 266b92faffb66af24d8ed2725beb80770a2d91f8#diff-0dd025320c7ecda2ea310ed7172d7f5a \| spark.mesos.role \| 1.5.0 \| SPARK-6284 \| d86bbb4e286f16f77ba125452b07827684eafeed#diff-02a6d899f7a529eb7cfbb12182a110b0 \| The following appears in the document \| \| \| \| spark.mesos.driverEnv.[EnvironmentVariableName] \| 2.1.0 \| SPARK-16194 \| 235cb256d06653bcde4c3ed6b081503a94996321#diff-b964c449b99c51f0a5fd77270b2951a4 \| spark.mesos.dispatcher.driverDefault.[PropertyName] \| 2.1.0 \| SPARK-16927 and SPARK-16923 \| eca58755fbbc11937b335ad953a3caff89b818e6#diff-b964c449b99c51f0a5fd77270b2951a4 \| ### Why are the changes needed? Supplemental configuration version information. ### Does this PR introduce any user-facing change? 'No'. ### How was this patch tested? Exists UT Closes #27863 from beliefer/add-version-to-mesos-config. Authored-by: beliefer <beliefer@163.com> Signed-off-by: HyukjinKwon <gurwls223@apache.org>	2020-03-12 11:02:29 +09:00
beliefer	1254c88034	[SPARK-31118][K8S][DOC] Add version information to the configuration of K8S ### What changes were proposed in this pull request? Add version information to the configuration of `K8S`. I sorted out some information show below. Item name \| Since version \| JIRA ID \| Commit ID \| Note -- \| -- \| -- \| -- \| -- spark.kubernetes.context \| 3.0.0 \| SPARK-25887 \| c542c247bbfe1214c0bf81076451718a9e8931dc#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.driver.master \| 3.0.0 \| SPARK-30371 \| f14061c6a4729ad419902193aa23575d8f17f597#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.namespace \| 2.3.0 \| SPARK-18278 \| e9b2070ab2d04993b1c0c1d6c6aba249e6664c8d#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.container.image \| 2.3.0 \| SPARK-22994 \| b94debd2b01b87ef1d2a34d48877e38ade0969e6#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.driver.container.image \| 2.3.0 \| SPARK-22807 \| fb3636b482be3d0940345b1528c1d5090bbc25e6#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.executor.container.image \| 2.3.0 \| SPARK-22807 \| fb3636b482be3d0940345b1528c1d5090bbc25e6#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.container.image.pullPolicy \| 2.3.0 \| SPARK-22807 \| fb3636b482be3d0940345b1528c1d5090bbc25e6#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.container.image.pullSecrets \| 2.4.0 \| SPARK-23668 \| cccaaa14ad775fb981e501452ba2cc06ff5c0f0a#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.submission.requestTimeout \| 3.0.0 \| SPARK-27023 \| e9e8bb33ef9ad785473ded168bc85867dad4ee70#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.submission.connectionTimeout \| 3.0.0 \| SPARK-27023 \| e9e8bb33ef9ad785473ded168bc85867dad4ee70#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.driver.requestTimeout \| 3.0.0 \| SPARK-27023 \| e9e8bb33ef9ad785473ded168bc85867dad4ee70#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.driver.connectionTimeout \| 3.0.0 \| SPARK-27023 \| e9e8bb33ef9ad785473ded168bc85867dad4ee70#diff-6e882d5561424e7e6651eb46f10104b8 \| KUBERNETES_AUTH_DRIVER_CONF_PREFIX.serviceAccountName \| 2.3.0 \| SPARK-18278 \| e9b2070ab2d04993b1c0c1d6c6aba249e6664c8d#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.authenticate.driver KUBERNETES_AUTH_EXECUTOR_CONF_PREFIX.serviceAccountName \| 3.1.0 \| SPARK-30122 \| f9f06eee9853ad4b6458ac9d31233e729a1ca226#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.authenticate.executor spark.kubernetes.driver.limit.cores \| 2.3.0 \| SPARK-22646 \| 3f4060c340d6bac412e8819c4388ccba226efcf3#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.driver.request.cores \| 3.0.0 \| SPARK-27754 \| 1a8c09334db87b0e938c38cd6b59d326bdcab3c3#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.submitInDriver \| 2.4.0 \| SPARK-22839 \| f15906da153f139b698e192ec6f82f078f896f1e#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.executor.limit.cores \| 2.3.0 \| SPARK-18278 \| e9b2070ab2d04993b1c0c1d6c6aba249e6664c8d#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.executor.scheduler.name \| 3.0.0 \| SPARK-29436 \| f800fa383131559c4e841bf062c9775d09190935#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.executor.request.cores \| 2.4.0 \| SPARK-23285 \| fe2b7a4568d65a62da6e6eb00fff05f248b4332c#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.driver.pod.name \| 2.3.0 \| SPARK-18278 \| e9b2070ab2d04993b1c0c1d6c6aba249e6664c8d#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.driver.resourceNamePrefix \| 3.0.0 \| SPARK-25876 \| 6be272b75b4ae3149869e19df193675cc4117763#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.executor.podNamePrefix \| 2.3.0 \| SPARK-18278 \| e9b2070ab2d04993b1c0c1d6c6aba249e6664c8d#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.allocation.batch.size \| 2.3.0 \| SPARK-18278 \| e9b2070ab2d04993b1c0c1d6c6aba249e6664c8d#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.allocation.batch.delay \| 2.3.0 \| SPARK-18278 \| e9b2070ab2d04993b1c0c1d6c6aba249e6664c8d#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.executor.lostCheck.maxAttempts \| 2.3.0 \| SPARK-18278 \| e9b2070ab2d04993b1c0c1d6c6aba249e6664c8d#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.submission.waitAppCompletion \| 2.3.0 \| SPARK-22646 \| 3f4060c340d6bac412e8819c4388ccba226efcf3#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.report.interval \| 2.3.0 \| SPARK-22646 \| 3f4060c340d6bac412e8819c4388ccba226efcf3#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.executor.apiPollingInterval \| 2.4.0 \| SPARK-24248 \| 270a9a3cac25f3e799460320d0fc94ccd7ecfaea#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.executor.eventProcessingInterval \| 2.4.0 \| SPARK-24248 \| 270a9a3cac25f3e799460320d0fc94ccd7ecfaea#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.memoryOverheadFactor \| 2.4.0 \| SPARK-23984 \| 1a644afbac35c204f9ad55f86999319a9ab458c6#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.pyspark.pythonVersion \| 2.4.0 \| SPARK-23984 \| a791c29bd824adadfb2d85594bc8dad4424df936#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.kerberos.krb5.path \| 3.0.0 \| SPARK-23257 \| 6c9c84ffb9c8d98ee2ece7ba4b010856591d383d#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.kerberos.krb5.configMapName \| 3.0.0 \| SPARK-23257 \| 6c9c84ffb9c8d98ee2ece7ba4b010856591d383d#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.hadoop.configMapName \| 3.0.0 \| SPARK-23257 \| 6c9c84ffb9c8d98ee2ece7ba4b010856591d383d#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.kerberos.tokenSecret.name \| 3.0.0 \| SPARK-23257 \| 6c9c84ffb9c8d98ee2ece7ba4b010856591d383d#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.kerberos.tokenSecret.itemKey \| 3.0.0 \| SPARK-23257 \| 6c9c84ffb9c8d98ee2ece7ba4b010856591d383d#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.resource.type \| 2.4.1 \| SPARK-25021 \| 9031c784847353051bc0978f63ef4146ae9095ff#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.local.dirs.tmpfs \| 3.0.0 \| SPARK-25262 \| da6fa3828bb824b65f50122a8a0a0d4741551257#diff-6e882d5561424e7e6651eb46f10104b8 \| It exists in branch-3.0, but in pom.xml it is 2.4.0-snapshot spark.kubernetes.driver.podTemplateFile \| 3.0.0 \| SPARK-24434 \| f6cc354d83c2c9a757f9b507aadd4dbdc5825cca#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.executor.podTemplateFile \| 3.0.0 \| SPARK-24434 \| f6cc354d83c2c9a757f9b507aadd4dbdc5825cca#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.driver.podTemplateContainerName \| 3.0.0 \| SPARK-24434 \| f6cc354d83c2c9a757f9b507aadd4dbdc5825cca#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.executor.podTemplateContainerName \| 3.0.0 \| SPARK-24434 \| f6cc354d83c2c9a757f9b507aadd4dbdc5825cca#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.executor.deleteOnTermination \| 3.0.0 \| SPARK-25515 \| 0c2935b01def8a5f631851999d9c2d57b63763e6#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.dynamicAllocation.deleteGracePeriod \| 3.0.0 \| SPARK-28487 \| 0343854f54b48b206ca434accec99355011560c2#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.appKillPodDeletionGracePeriod \| 3.0.0 \| SPARK-24793 \| 05168e725d2a17c4164ee5f9aa068801ec2454f4#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.file.upload.path \| 3.0.0 \| SPARK-23153 \| 5e74570c8f5e7dfc1ca1c53c177827c5cea57bf1#diff-6e882d5561424e7e6651eb46f10104b8 \| The following appears in the document \| \| \| \| spark.kubernetes.authenticate.submission.caCertFile \| 2.3.0 \| SPARK-22646 \| 3f4060c340d6bac412e8819c4388ccba226efcf3#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.authenticate.submission.clientKeyFile \| 2.3.0 \| SPARK-22646 \| 3f4060c340d6bac412e8819c4388ccba226efcf3#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.authenticate.submission.clientCertFile \| 2.3.0 \| SPARK-22646 \| 3f4060c340d6bac412e8819c4388ccba226efcf3#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.authenticate.submission.oauthToken \| 2.3.0 \| SPARK-22646 \| 3f4060c340d6bac412e8819c4388ccba226efcf3#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.authenticate.submission.oauthTokenFile \| 2.3.0 \| SPARK-22646 \| 3f4060c340d6bac412e8819c4388ccba226efcf3#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.authenticate.driver.caCertFile \| 2.3.0 \| SPARK-18278 \| e9b2070ab2d04993b1c0c1d6c6aba249e6664c8d#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.authenticate.driver.clientKeyFile \| 2.3.0 \| SPARK-18278 \| e9b2070ab2d04993b1c0c1d6c6aba249e6664c8d#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.authenticate.driver.clientCertFile \| 2.3.0 \| SPARK-18278 \| e9b2070ab2d04993b1c0c1d6c6aba249e6664c8d#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.authenticate.driver.oauthToken \| 2.3.0 \| SPARK-18278 \| e9b2070ab2d04993b1c0c1d6c6aba249e6664c8d#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.authenticate.driver.oauthTokenFile \| 2.3.0 \| SPARK-18278 \| e9b2070ab2d04993b1c0c1d6c6aba249e6664c8d#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.authenticate.driver.mounted.caCertFile \| 2.3.0 \| SPARK-18278 \| e9b2070ab2d04993b1c0c1d6c6aba249e6664c8d#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.authenticate.driver.mounted.clientKeyFile \| 2.3.0 \| SPARK-18278 \| e9b2070ab2d04993b1c0c1d6c6aba249e6664c8d#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.authenticate.driver.mounted.clientCertFile \| 2.3.0 \| SPARK-18278 \| e9b2070ab2d04993b1c0c1d6c6aba249e6664c8d#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.authenticate.driver.mounted.oauthTokenFile \| 2.3.0 \| SPARK-18278 \| e9b2070ab2d04993b1c0c1d6c6aba249e6664c8d#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.authenticate.caCertFile \| 2.4.0 \| SPARK-23146 \| 571a6f0574e50e53cea403624ec3795cd03aa204#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.authenticate.clientKeyFile \| 2.4.0 \| SPARK-23146 \| 571a6f0574e50e53cea403624ec3795cd03aa204#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.authenticate.clientCertFile \| 2.4.0 \| SPARK-23146 \| 571a6f0574e50e53cea403624ec3795cd03aa204#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.authenticate.oauthToken \| 2.4.0 \| SPARK-23146 \| 571a6f0574e50e53cea403624ec3795cd03aa204#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.authenticate.oauthTokenFile \| 2.4.0 \| SPARK-23146 \| 571a6f0574e50e53cea403624ec3795cd03aa204#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.driver.label.[LabelName] \| 2.3.0 \| SPARK-22646 \| 3f4060c340d6bac412e8819c4388ccba226efcf3#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.driver.annotation.[AnnotationName] \| 2.3.0 \| SPARK-22646 \| 3f4060c340d6bac412e8819c4388ccba226efcf3#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.executor.label.[LabelName] \| 2.3.0 \| SPARK-22646 \| 3f4060c340d6bac412e8819c4388ccba226efcf3#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.executor.annotation.[AnnotationName] \| 2.3.0 \| SPARK-22646 \| 3f4060c340d6bac412e8819c4388ccba226efcf3#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.node.selector.[labelKey] \| 2.3.0 \| SPARK-18278 \| e9b2070ab2d04993b1c0c1d6c6aba249e6664c8d#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.driverEnv.[EnvironmentVariableName] \| 2.3.0 \| SPARK-22646 \| 3f4060c340d6bac412e8819c4388ccba226efcf3#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.driver.secrets.[SecretName] \| 2.3.0 \| SPARK-22757 \| 171f6ddadc6185ffcc6ad82e5f48952fb49095b2#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.executor.secrets.[SecretName] \| 2.3.0 \| SPARK-22757 \| 171f6ddadc6185ffcc6ad82e5f48952fb49095b2#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.driver.secretKeyRef.[EnvName] \| 2.4.0 \| SPARK-24232 \| 21e1fc7d4aed688d7b685be6ce93f76752159c98#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.executor.secretKeyRef.[EnvName] \| 2.4.0 \| SPARK-24232 \| 21e1fc7d4aed688d7b685be6ce93f76752159c98#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.driver.volumes.[VolumeType].[VolumeName].mount.path \| 2.4.0 \| SPARK-23529 \| 5ff1b9ba1983d5601add62aef64a3e87d07050eb#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.driver.volumes.[VolumeType].[VolumeName].mount.subPath \| 3.0.0 \| SPARK-25960 \| 3df307aa515b3564686e75d1b71754bbcaaf2dec#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.driver.volumes.[VolumeType].[VolumeName].mount.readOnly \| 2.4.0 \| SPARK-23529 \| 5ff1b9ba1983d5601add62aef64a3e87d07050eb#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.driver.volumes.[VolumeType].[VolumeName].options.[OptionName] \| 2.4.0 \| SPARK-23529 \| 5ff1b9ba1983d5601add62aef64a3e87d07050eb#diff-b5527f236b253e0d9f5db5164bdb43e9 \| spark.kubernetes.executor.volumes.[VolumeType].[VolumeName].mount.path \| 2.4.0 \| SPARK-23529 \| 5ff1b9ba1983d5601add62aef64a3e87d07050eb#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.executor.volumes.[VolumeType].[VolumeName].mount.subPath \| 3.0.0 \| SPARK-25960 \| 3df307aa515b3564686e75d1b71754bbcaaf2dec#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.executor.volumes.[VolumeType].[VolumeName].mount.readOnly \| 2.4.0 \| SPARK-23529 \| 5ff1b9ba1983d5601add62aef64a3e87d07050eb#diff-6e882d5561424e7e6651eb46f10104b8 \| spark.kubernetes.executor.volumes.[VolumeType].[VolumeName].options.[OptionName] \| 2.4.0 \| SPARK-23529 \| 5ff1b9ba1983d5601add62aef64a3e87d07050eb#diff-b5527f236b253e0d9f5db5164bdb43e9 \| ### Why are the changes needed? Supplemental configuration version information. ### Does this PR introduce any user-facing change? 'No' ### How was this patch tested? Exists UT Closes #27875 from beliefer/add-version-to-k8s-config. Authored-by: beliefer <beliefer@163.com> Signed-off-by: HyukjinKwon <gurwls223@apache.org>	2020-03-12 09:54:08 +09:00
beliefer	0722dc5fb8	[SPARK-31092][YARN][DOC] Add version information to the configuration of Yarn ### What changes were proposed in this pull request? Add version information to the configuration of `Yarn`. I sorted out some information show below. Item name \| Since version \| JIRA ID \| Commit ID \| Note -- \| -- \| -- \| -- \| -- spark.yarn.tags \| 1.5.0 \| SPARK-9782 \| 9b731fad2b43ca18f3c5274062d4c7bc2622ab72#diff-b050df3f55b82065803d6e83453b9706 \| spark.yarn.priority \| 3.0.0 \| SPARK-29603 \| 4615769736f4c052ae1a2de26e715e229154cd2f#diff-4804e0f83ca7f891183eb0db229b4b9a \| spark.yarn.am.attemptFailuresValidityInterval \| 1.6.0 \| SPARK-10739 \| f97e9323b526b3d0b0fee0ca03f4276f37bb5750#diff-b050df3f55b82065803d6e83453b9706 \| spark.yarn.executor.failuresValidityInterval \| 2.0.0 \| SPARK-6735 \| 8b44bd52fa40c0fc7d34798c3654e31533fd3008#diff-14b8ed2ef4e3da985300b8d796a38fa9 \| spark.yarn.maxAppAttempts \| 1.3.0 \| SPARK-2165 \| 8fdd48959c93b9cf809f03549e2ae6c4687d1fcd#diff-b050df3f55b82065803d6e83453b9706 \| spark.yarn.user.classpath.first \| 1.3.0 \| SPARK-5087 \| 8d45834debc6986e61831d0d6e982d5528dccc51#diff-b050df3f55b82065803d6e83453b9706 \| spark.yarn.config.gatewayPath \| 1.5.0 \| SPARK-8302 \| 37bf76a2de2143ec6348a3d43b782227849520cc#diff-b050df3f55b82065803d6e83453b9706 \| spark.yarn.config.replacementPath \| 1.5.0 \| SPARK-8302 \| 37bf76a2de2143ec6348a3d43b782227849520cc#diff-b050df3f55b82065803d6e83453b9706 \| spark.yarn.queue \| 1.0.0 \| SPARK-1126 \| 1617816090e7b20124a512a43860a21232ebf511#diff-ae6a41a938a767e5bb97b5d738371a5b \| spark.yarn.historyServer.address \| 1.0.0 \| SPARK-1408 \| 0058b5d2c74147d24b127a5432f89ebc7050dc18#diff-923ae58523a12397f74dd590744b8b41 \| spark.yarn.historyServer.allowTracking \| 2.2.0 \| SPARK-19554 \| 4661d30b988bf773ab45a15b143efb2908d33743#diff-4804e0f83ca7f891183eb0db229b4b9a \| spark.yarn.archive \| 2.0.0 \| SPARK-13577 \| 07f1c5447753a3d593cd6ececfcb03c11b1cf8ff#diff-14b8ed2ef4e3da985300b8d796a38fa9 \| spark.yarn.jars \| 2.0.0 \| SPARK-13577 \| 07f1c5447753a3d593cd6ececfcb03c11b1cf8ff#diff-14b8ed2ef4e3da985300b8d796a38fa9 \| spark.yarn.dist.archives \| 1.0.0 \| SPARK-1126 \| 1617816090e7b20124a512a43860a21232ebf511#diff-ae6a41a938a767e5bb97b5d738371a5b \| spark.yarn.dist.files \| 1.0.0 \| SPARK-1126 \| 1617816090e7b20124a512a43860a21232ebf511#diff-ae6a41a938a767e5bb97b5d738371a5b \| spark.yarn.dist.jars \| 2.0.0 \| SPARK-12343 \| 8ba2b7f28fee39c4839e5ea125bd25f5091a3a1e#diff-14b8ed2ef4e3da985300b8d796a38fa9 \| spark.yarn.preserve.staging.files \| 1.1.0 \| SPARK-2933 \| b92d823ad13f6fcc325eeb99563bea543871c6aa#diff-85a1f4b2810b3e11b8434dcefac5bb85 \| spark.yarn.submit.file.replication \| 0.8.1 \| None \| 4668fcb9ff8f9c176c4866480d52dde5d67c8522#diff-b050df3f55b82065803d6e83453b9706 \| spark.yarn.submit.waitAppCompletion \| 1.4.0 \| SPARK-3591 \| b65bad65c3500475b974ca0219f218eef296db2c#diff-b050df3f55b82065803d6e83453b9706 \| spark.yarn.report.interval \| 0.9.0 \| None \| ebdfa6bb9766209bc5a3c4241fa47141c5e9c5cb#diff-e0a7ae95b6d8e04a67ebca0945d27b65 \| spark.yarn.clientLaunchMonitorInterval \| 2.3.0 \| SPARK-16019 \| 1cad31f00644d899d8e74d58c6eb4e9f72065473#diff-4804e0f83ca7f891183eb0db229b4b9a \| spark.yarn.am.waitTime \| 1.3.0 \| SPARK-3779 \| 253b72b56fe908bbab5d621eae8a5f359c639dfd#diff-87125050a2e2eaf87ea83aac9c19b200 \| spark.yarn.metrics.namespace \| 2.4.0 \| SPARK-24594 \| d2436a85294a178398525c37833dae79d45c1452#diff-4804e0f83ca7f891183eb0db229b4b9a \| spark.yarn.am.nodeLabelExpression \| 1.6.0 \| SPARK-7173 \| 7db3610327d0725ec2ad378bc873b127a59bb87a#diff-b050df3f55b82065803d6e83453b9706 \| spark.yarn.containerLauncherMaxThreads \| 1.2.0 \| SPARK-1713 \| 1f4a648d4e30e837d6cf3ea8de1808e2254ad70b#diff-801a04f9e67321f3203399f7f59234c1 \| spark.yarn.max.executor.failures \| 1.0.0 \| SPARK-1183 \| 698373211ef3cdf841c82d48168cd5dbe00a57b4#diff-0c239e58b37779967e0841fb42f3415a \| spark.yarn.scheduler.reporterThread.maxFailures \| 1.2.0 \| SPARK-3304 \| 11c10df825419372df61a8d23c51e8c3cc78047f#diff-85a1f4b2810b3e11b8434dcefac5bb85 \| spark.yarn.scheduler.heartbeat.interval-ms \| 0.8.1 \| None \| ee22be0e6c302fb2cdb24f83365c2b8a43a1baab#diff-87125050a2e2eaf87ea83aac9c19b200 \| spark.yarn.scheduler.initial-allocation.interval \| 1.4.0 \| SPARK-7533 \| 3ddf051ee7256f642f8a17768d161c7b5f55c7e1#diff-87125050a2e2eaf87ea83aac9c19b200 \| spark.yarn.am.finalMessageLimit \| 2.4.0 \| SPARK-25174 \| f8346d2fc01f1e881e4e3f9c4499bf5f9e3ceb3f#diff-4804e0f83ca7f891183eb0db229b4b9a \| spark.yarn.am.cores \| 1.3.0 \| SPARK-1507 \| 2be82b1e66cd188456bbf1e5abb13af04d1629d5#diff-746d34aa06bfa57adb9289011e725472 \| spark.yarn.am.extraJavaOptions \| 1.3.0 \| SPARK-5087 \| 8d45834debc6986e61831d0d6e982d5528dccc51#diff-b050df3f55b82065803d6e83453b9706 \| spark.yarn.am.extraLibraryPath \| 1.4.0 \| SPARK-7281 \| 7b5dd3e3c0030087eea5a8224789352c03717c1d#diff-b050df3f55b82065803d6e83453b9706 \| spark.yarn.am.memoryOverhead \| 1.3.0 \| SPARK-1953 \| e96645206006a009e5c1a23bbd177dcaf3ef9b83#diff-746d34aa06bfa57adb9289011e725472 \| spark.yarn.am.memory \| 1.3.0 \| SPARK-1953 \| e96645206006a009e5c1a23bbd177dcaf3ef9b83#diff-746d34aa06bfa57adb9289011e725472 \| spark.driver.appUIAddress \| 1.1.0 \| SPARK-1291 \| 72ea56da8e383c61c6f18eeefef03b9af00f5158#diff-2b4617e158e9c5999733759550440b96 \| spark.yarn.executor.nodeLabelExpression \| 1.4.0 \| SPARK-6470 \| 82fee9d9aad2c9ba2fb4bd658579fe99218cafac#diff-d4620cf162e045960d84c88b2e0aa428 \| spark.yarn.unmanagedAM.enabled \| 3.0.0 \| SPARK-22404 \| f06bc0cd1dee2a58e04ebf24bf719a2f7ef2dc4e#diff-4804e0f83ca7f891183eb0db229b4b9a \| spark.yarn.rolledLog.includePattern \| 2.0.0 \| SPARK-15990 \| 272a2f78f3ff801b94a81fa8fcc6633190eaa2f4#diff-14b8ed2ef4e3da985300b8d796a38fa9 \| spark.yarn.rolledLog.excludePattern \| 2.0.0 \| SPARK-15990 \| 272a2f78f3ff801b94a81fa8fcc6633190eaa2f4#diff-14b8ed2ef4e3da985300b8d796a38fa9 \| spark.yarn.user.jar \| 1.1.0 \| SPARK-1395 \| e380767de344fd6898429de43da592658fd86a39#diff-50e237ea17ce94c3ccfc44143518a5f7 \| spark.yarn.secondary.jars \| 0.9.2 \| SPARK-1870 \| 1d3aab96120c6770399e78a72b5692cf8f61a144#diff-50b743cff4885220c828b16c44eeecfd \| spark.yarn.cache.filenames \| 2.0.0 \| SPARK-14602 \| f47dbf27fa034629fab12d0f3c89ab75edb03f86#diff-14b8ed2ef4e3da985300b8d796a38fa9 \| spark.yarn.cache.sizes \| 2.0.0 \| SPARK-14602 \| f47dbf27fa034629fab12d0f3c89ab75edb03f86#diff-14b8ed2ef4e3da985300b8d796a38fa9 \| spark.yarn.cache.timestamps \| 2.0.0 \| SPARK-14602 \| f47dbf27fa034629fab12d0f3c89ab75edb03f86#diff-14b8ed2ef4e3da985300b8d796a38fa9 \| spark.yarn.cache.visibilities \| 2.0.0 \| SPARK-14602 \| f47dbf27fa034629fab12d0f3c89ab75edb03f86#diff-14b8ed2ef4e3da985300b8d796a38fa9 \| spark.yarn.cache.types \| 2.0.0 \| SPARK-14602 \| f47dbf27fa034629fab12d0f3c89ab75edb03f86#diff-14b8ed2ef4e3da985300b8d796a38fa9 \| spark.yarn.cache.confArchive \| 2.0.0 \| SPARK-14602 \| f47dbf27fa034629fab12d0f3c89ab75edb03f86#diff-14b8ed2ef4e3da985300b8d796a38fa9 \| spark.yarn.blacklist.executor.launch.blacklisting.enabled \| 2.4.0 \| SPARK-16630 \| b56e9c613fb345472da3db1a567ee129621f6bf3#diff-4804e0f83ca7f891183eb0db229b4b9a \| spark.yarn.exclude.nodes \| 3.0.0 \| SPARK-26688 \| caceaec93203edaea1d521b88e82ef67094cdea9#diff-4804e0f83ca7f891183eb0db229b4b9a \| The following appears in the document \| \| \| \| spark.yarn.am.resource.{resource-type}.amount \| 3.0.0 \| SPARK-20327 \| 3946de773498621f88009c309254b019848ed490#diff-4804e0f83ca7f891183eb0db229b4b9a \| spark.yarn.driver.resource.{resource-type}.amount \| 3.0.0 \| SPARK-20327 \| 3946de773498621f88009c309254b019848ed490#diff-4804e0f83ca7f891183eb0db229b4b9a \| spark.yarn.executor.resource.{resource-type}.amount \| 3.0.0 \| SPARK-20327 \| 3946de773498621f88009c309254b019848ed490#diff-4804e0f83ca7f891183eb0db229b4b9a \| spark.yarn.appMasterEnv.[EnvironmentVariableName] \| 1.1.0 \| SPARK-1680 \| 7b798e10e214cd407d3399e2cab9e3789f9a929e#diff-50e237ea17ce94c3ccfc44143518a5f7 \| spark.yarn.kerberos.relogin.period \| 2.3.0 \| SPARK-22290 \| dc2714da50ecba1bf1fdf555a82a4314f763a76e#diff-4804e0f83ca7f891183eb0db229b4b9a \| ### Why are the changes needed? Supplemental configuration version information. ### Does this PR introduce any user-facing change? 'No'. ### How was this patch tested? Exists UT Closes #27856 from beliefer/add-version-to-yarn-config. Authored-by: beliefer <beliefer@163.com> Signed-off-by: HyukjinKwon <gurwls223@apache.org>	2020-03-12 09:52:57 +09:00
beliefer	c1b2675f2e	[SPARK-31002][CORE][DOC][FOLLOWUP] Add version information to the configuration of Core ### What changes were proposed in this pull request? This PR follows up https://github.com/apache/spark/pull/27847. I sorted out some information show below. Item name \| Since version \| JIRA ID \| Commit ID \| Note -- \| -- \| -- \| -- \| -- spark.yarn.isPython \| 1.5.0 \| SPARK-5479 \| 38112905bc3b33f2ae75274afba1c30e116f6e46#diff-4d2ab44195558d5a9d5f15b8803ef39d \| spark.task.cpus \| 0.5.0 \| None \| e5c4cd8a5e188592f8786a265c0cd073c69ac886#diff-391214d132a0fb4478f4f9c2313d8966 \| spark.dynamicAllocation.enabled \| 1.2.0 \| SPARK-3795 \| 8d59b37b02eb36f37bcefafb952519d7dca744ad#diff-364713d7776956cb8b0a771e9b62f82d \| spark.dynamicAllocation.testing \| 1.2.0 \| SPARK-3795 \| 8d59b37b02eb36f37bcefafb952519d7dca744ad#diff-364713d7776956cb8b0a771e9b62f82d \| spark.dynamicAllocation.minExecutors \| 1.2.0 \| SPARK-3795 \| 8d59b37b02eb36f37bcefafb952519d7dca744ad#diff-364713d7776956cb8b0a771e9b62f82d \| spark.dynamicAllocation.initialExecutors \| 1.3.0 \| SPARK-4585 \| b2047b55c5fc85de6b63276d8ab9610d2496e08b#diff-b096353602813e47074ace09a3890d56 \| spark.dynamicAllocation.maxExecutors \| 1.2.0 \| SPARK-3795 \| 8d59b37b02eb36f37bcefafb952519d7dca744ad#diff-364713d7776956cb8b0a771e9b62f82d \| spark.dynamicAllocation.executorAllocationRatio \| 2.4.0 \| SPARK-22683 \| 55c4ca88a3b093ee197a8689631be8d1fac1f10f#diff-6bdad48cfc34314e89599655442ff210 \| spark.dynamicAllocation.cachedExecutorIdleTimeout \| 1.4.0 \| SPARK-7955 \| 6faaf15ba311bc3a79aae40a6c9c4befabb6889f#diff-b096353602813e47074ace09a3890d56 \| spark.dynamicAllocation.executorIdleTimeout \| 1.2.0 \| SPARK-3795 \| 8d59b37b02eb36f37bcefafb952519d7dca744ad#diff-364713d7776956cb8b0a771e9b62f82d \| spark.dynamicAllocation.shuffleTracking.enabled \| 3.0.0 \| SPARK-27963 \| 2ddeff97d7329942a98ef363991eeabc3fa71a76#diff-6bdad48cfc34314e89599655442ff210 \| spark.dynamicAllocation.shuffleTimeout \| 3.0.0 \| SPARK-27963 \| 2ddeff97d7329942a98ef363991eeabc3fa71a76#diff-6bdad48cfc34314e89599655442ff210 \| spark.dynamicAllocation.schedulerBacklogTimeout \| 1.2.0 \| SPARK-3795 \| 8d59b37b02eb36f37bcefafb952519d7dca744ad#diff-364713d7776956cb8b0a771e9b62f82d \| spark.dynamicAllocation.sustainedSchedulerBacklogTimeout \| 1.2.0 \| SPARK-3795 \| 8d59b37b02eb36f37bcefafb952519d7dca744ad#diff-364713d7776956cb8b0a771e9b62f82d \| spark.locality.wait \| 0.5.0 \| None \| e5c4cd8a5e188592f8786a265c0cd073c69ac886#diff-391214d132a0fb4478f4f9c2313d8966 \| spark.shuffle.service.enabled \| 1.2.0 \| SPARK-3796 \| f55218aeb1e9d638df6229b36a59a15ce5363482#diff-2b643ea78c1add0381754b1f47eec132 \| Constants.SHUFFLE_SERVICE_FETCH_RDD_ENABLED \| 3.0.0 \| SPARK-27677 \| e9f3f62b2c0f521f3cc23fef381fc6754853ad4f#diff-6bdad48cfc34314e89599655442ff210 \| spark.shuffle.service.fetch.rdd.enabled spark.shuffle.service.db.enabled \| 3.0.0 \| SPARK-26288 \| 8b0aa59218c209d39cbba5959302d8668b885cf6#diff-6bdad48cfc34314e89599655442ff210 \| spark.shuffle.service.port \| 1.2.0 \| SPARK-3796 \| f55218aeb1e9d638df6229b36a59a15ce5363482#diff-2b643ea78c1add0381754b1f47eec132 \| spark.kerberos.keytab \| 3.0.0 \| SPARK-25372 \| 51540c2fa677658be954c820bc18ba748e4c8583#diff-6bdad48cfc34314e89599655442ff210 \| spark.kerberos.principal \| 3.0.0 \| SPARK-25372 \| 51540c2fa677658be954c820bc18ba748e4c8583#diff-6bdad48cfc34314e89599655442ff210 \| spark.kerberos.relogin.period \| 3.0.0 \| SPARK-23781 \| 68dde3481ea458b0b8deeec2f99233c2d4c1e056#diff-6bdad48cfc34314e89599655442ff210 \| spark.kerberos.renewal.credentials \| 3.0.0 \| SPARK-26595 \| 2a67dbfbd341af166b1c85904875f26a6dea5ba8#diff-6bdad48cfc34314e89599655442ff210 \| spark.kerberos.access.hadoopFileSystems \| 3.0.0 \| SPARK-26766 \| d0443a74d185ec72b747fa39994fa9a40ce974cf#diff-6bdad48cfc34314e89599655442ff210 \| spark.executor.instances \| 1.0.0 \| SPARK-1126 \| 1617816090e7b20124a512a43860a21232ebf511#diff-4d2ab44195558d5a9d5f15b8803ef39d \| spark.yarn.dist.pyFiles \| 2.2.1 \| SPARK-21714 \| d10c9dc3f631a26dbbbd8f5c601ca2001a5d7c80#diff-6bdad48cfc34314e89599655442ff210 \| spark.task.maxDirectResultSize \| 2.0.0 \| SPARK-13830 \| 2ef4c5963bff3574fe17e669d703b25ddd064e5d#diff-5a0de266c82b95adb47d9bca714e1f1b \| spark.task.maxFailures \| 0.8.0 \| None \| 46eecd110a4017ea0c86cbb1010d0ccd6a5eb2ef#diff-264da78fe625d594eae59d1adabc8ae9 \| spark.task.reaper.enabled \| 2.0.3 \| SPARK-18761 \| 678d91c1d2283d9965a39656af9d383bad093ba8#diff-5a0de266c82b95adb47d9bca714e1f1b \| spark.task.reaper.killTimeout \| 2.0.3 \| SPARK-18761 \| 678d91c1d2283d9965a39656af9d383bad093ba8#diff-5a0de266c82b95adb47d9bca714e1f1b \| spark.task.reaper.pollingInterval \| 2.0.3 \| SPARK-18761 \| 678d91c1d2283d9965a39656af9d383bad093ba8#diff-5a0de266c82b95adb47d9bca714e1f1b \| spark.task.reaper.threadDump \| 2.0.3 \| SPARK-18761 \| 678d91c1d2283d9965a39656af9d383bad093ba8#diff-5a0de266c82b95adb47d9bca714e1f1b \| spark.blacklist.enabled \| 2.1.0 \| SPARK-17675 \| 9ce7d3e542e786c62f047c13f3001e178f76e06a#diff-6bdad48cfc34314e89599655442ff210 \| spark.blacklist.task.maxTaskAttemptsPerExecutor \| 2.1.0 \| SPARK-17675 \| 9ce7d3e542e786c62f047c13f3001e178f76e06a#diff-6bdad48cfc34314e89599655442ff210 \| spark.blacklist.task.maxTaskAttemptsPerNode \| 2.1.0 \| SPARK-17675 \| 9ce7d3e542e786c62f047c13f3001e178f76e06a#diff-6bdad48cfc34314e89599655442ff210 \| spark.blacklist.application.maxFailedTasksPerExecutor \| 2.2.0 \| SPARK-8425 \| 93cdb8a7d0f124b4db069fd8242207c82e263c52#diff-6bdad48cfc34314e89599655442ff210 \| spark.blacklist.stage.maxFailedTasksPerExecutor \| 2.1.0 \| SPARK-17675 \| 9ce7d3e542e786c62f047c13f3001e178f76e06a#diff-6bdad48cfc34314e89599655442ff210 \| spark.blacklist.application.maxFailedExecutorsPerNode \| 2.2.0 \| SPARK-8425 \| 93cdb8a7d0f124b4db069fd8242207c82e263c52#diff-6bdad48cfc34314e89599655442ff210 \| spark.blacklist.stage.maxFailedExecutorsPerNode \| 2.1.0 \| SPARK-17675 \| 9ce7d3e542e786c62f047c13f3001e178f76e06a#diff-6bdad48cfc34314e89599655442ff210 \| spark.blacklist.timeout \| 2.1.0 \| SPARK-17675 \| 9ce7d3e542e786c62f047c13f3001e178f76e06a#diff-6bdad48cfc34314e89599655442ff210 \| spark.blacklist.killBlacklistedExecutors \| 2.2.0 \| SPARK-16554 \| 6287c94f08200d548df5cc0a401b73b84f9968c4#diff-6bdad48cfc34314e89599655442ff210 \| spark.scheduler.executorTaskBlacklistTime \| 1.0.0 \| None \| ab747d39ddc7c8a314ed2fb26548fc5652af0d74#diff-bad3987c83bd22d46416d3dd9d208e76 \| spark.blacklist.application.fetchFailure.enabled \| 2.3.0 \| SPARK-13669 and SPARK-20898 \| 9e50a1d37a4cf0c34e20a7c1a910ceaff41535a2#diff-6bdad48cfc34314e89599655442ff210 \| spark.files.fetchFailure.unRegisterOutputOnHost \| 2.3.0 \| SPARK-19753 \| dccc0aa3cf957c8eceac598ac81ac82f03b52105#diff-6bdad48cfc34314e89599655442ff210 \| spark.scheduler.listenerbus.eventqueue.capacity \| 2.3.0 \| SPARK-20887 \| 629f38e171409da614fd635bd8dd951b7fde17a4#diff-6bdad48cfc34314e89599655442ff210 \| spark.scheduler.listenerbus.metrics.maxListenerClassesTimed \| 2.3.0 \| SPARK-20863 \| 2a23cdd078a7409d0bb92cf27718995766c41b1d#diff-6bdad48cfc34314e89599655442ff210 \| spark.scheduler.listenerbus.logSlowEvent \| 3.0.0 \| SPARK-30812 \| 68d7edf9497bea2f73707d32ab55dd8e53088e7c#diff-6bdad48cfc34314e89599655442ff210 \| spark.scheduler.listenerbus.logSlowEvent.threshold \| 3.0.0 \| SPARK-29001 \| 0346afa8fc348aa1b3f5110df747a64e3b2da388#diff-6bdad48cfc34314e89599655442ff210 \| ### Why are the changes needed? Supplemental configuration version information. ### Does this PR introduce any user-facing change? No ### How was this patch tested? Exists UT Closes #27852 from beliefer/add-version-to-core-config-part-two. Authored-by: beliefer <beliefer@163.com> Signed-off-by: HyukjinKwon <gurwls223@apache.org>	2020-03-12 09:52:20 +09:00
Wenchen Fan	0f0ccdadb1	[SPARK-31110][DOCS][SQL] refine sql doc for SELECT ### What changes were proposed in this pull request? A few improvements to the sql ref SELECT doc: 1. correct the syntax of SELECT query 2. correct the default of null sort order 3. correct the GROUP BY syntax 4. several minor fixes ### Why are the changes needed? refine document ### Does this PR introduce any user-facing change? N/A ### How was this patch tested? N/A Closes #27866 from cloud-fan/doc. Authored-by: Wenchen Fan <wenchen@databricks.com> Signed-off-by: Dongjoon Hyun <dongjoon@apache.org>	2020-03-11 16:52:40 -07:00
Holden Karau	2825237448	[SPARK-31062][K8S][TESTS] Improve spark decommissioning k8s test reliability ### What changes were proposed in this pull request? Replace a sleep with waiting for the first collect to happen to try and make the K8s test code more reliable. ### Why are the changes needed? Currently the Decommissioning test appears to be flaky in Jenkins. ### Does this PR introduce any user-facing change? No ### How was this patch tested? Ran K8s test suite in a loop on minikube on my desktop for 10 iterations without this test failing on any of the runs. Closes #27858 from holdenk/SPARK-31062-Improve-Spark-Decommissioning-K8s-test-teliability. Authored-by: Holden Karau <hkarau@apple.com> Signed-off-by: Holden Karau <hkarau@apple.com>	2020-03-11 14:42:31 -07:00
Huaxin Gao	a1a665bece	[SPARK-31077][ML] Remove ChiSqSelector dependency on mllib.ChiSqSelectorModel ### What changes were proposed in this pull request? ```ChiSqSelector ``` depends on ```mllib.ChiSqSelectorModel``` to do the selection logic. Will remove the dependency in this PR. ### Why are the changes needed? This PR is an intermediate PR. Removing ```ChiSqSelector``` dependency on ```mllib.ChiSqSelectorModel```. Next subtask will extract the common code between ```ChiSqSelector``` and ```FValueSelector``` and put in an abstract ```Selector```. ### Does this PR introduce any user-facing change? No ### How was this patch tested? New and existing tests Closes #27841 from huaxingao/chisq. Authored-by: Huaxin Gao <huaxing@us.ibm.com> Signed-off-by: Sean Owen <srowen@gmail.com>	2020-03-11 13:51:49 -05:00
Wenchen Fan	8efb71013d	[SPARK-31091] Revert SPARK-24640 Return `NULL` from `size(NULL)` by default ### What changes were proposed in this pull request? This PR reverts https://github.com/apache/spark/pull/26051 and https://github.com/apache/spark/pull/26066 ### Why are the changes needed? There is no standard requiring that `size(null)` must return null, and returning -1 looks reasonable as well. This is kind of a cosmetic change and we should avoid it if it breaks existing queries. This is similar to reverting TRIM function parameter order change. ### Does this PR introduce any user-facing change? Yes, change the behavior of `size(null)` back to be the same as 2.4. ### How was this patch tested? N/A Closes #27834 from cloud-fan/revert. Authored-by: Wenchen Fan <wenchen@databricks.com> Signed-off-by: Dongjoon Hyun <dongjoon@apache.org>	2020-03-11 09:55:24 -07:00
Wenchen Fan	5be0d04f16	[SPARK-31117][SQL][TEST] reduce the test time of DateTimeUtilsSuite ### What changes were proposed in this pull request? `DateTimeUtilsSuite.daysToMicros and microsToDays` takes 30 seconds, which is too long for a UT. This PR changes the test to check random data, to reduce testing time. Now this test takes 1 second. ### Why are the changes needed? make test faster ### Does this PR introduce any user-facing change? no ### How was this patch tested? N/A Closes #27873 from cloud-fan/test. Authored-by: Wenchen Fan <wenchen@databricks.com> Signed-off-by: Wenchen Fan <wenchen@databricks.com>	2020-03-11 23:47:13 +08:00
Nicholas Chammas	3341021add	[SPARK-31041][BUILD] Show Maven errors from within make-distribution.sh ### What changes were proposed in this pull request? This PR makes `dev/make-distribution.sh` a bit easier to use by not hiding errors thrown by Maven. As a supporting change, this PR also suppresses progress bar output from curl and wget that may be output from within `build/mvn`. ### Why are the changes needed? It's surprising for command-line options to be position-dependent. The errors that get thrown when you pass the correct option (like `--pip`) but in the wrong order are confusing. ### Does this PR introduce any user-facing change? No. ### How was this patch tested? I ran a few invocations of `make-distribution.sh` to confirm that, when passed incorrect options, I can actually see the errors thrown directly by Maven. Closes #27800 from nchammas/SPARK-31041-make-distribution. Authored-by: Nicholas Chammas <nicholas.chammas@liveramp.com> Signed-off-by: Sean Owen <srowen@gmail.com>	2020-03-11 08:22:02 -05:00
Maxim Gekk	3d3e366aa8	[SPARK-31076][SQL] Convert Catalyst's DATE/TIMESTAMP to Java Date/Timestamp via local date-time ### What changes were proposed in this pull request? In the PR, I propose to change conversion of java.sql.Timestamp/Date values to/from internal values of Catalyst's TimestampType/DateType before cutover day `1582-10-15` of Gregorian calendar. I propose to construct local date-time from microseconds/days since the epoch. Take each date-time component `year`, `month`, `day`, `hour`, `minute`, `second` and `second fraction`, and construct java.sql.Timestamp/Date using the extracted components. ### Why are the changes needed? This will rebase underlying time/date offset in the way that collected java.sql.Timestamp/Date values will have the same local time-date component as the original values in Gregorian calendar. Here is the example which demonstrates the issue: ```sql scala> sql("select date '1100-10-10'").collect() res1: Array[org.apache.spark.sql.Row] = Array([1100-10-03]) ``` ### Does this PR introduce any user-facing change? Yes, after the changes: ```sql scala> sql("select date '1100-10-10'").collect() res0: Array[org.apache.spark.sql.Row] = Array([1100-10-10]) ``` ### How was this patch tested? By running `DateTimeUtilsSuite`, `DateFunctionsSuite` and `DateExpressionsSuite`. Closes #27807 from MaxGekk/rebase-timestamp-before-1582. Authored-by: Maxim Gekk <max.gekk@gmail.com> Signed-off-by: Wenchen Fan <wenchen@databricks.com>	2020-03-11 20:53:56 +08:00
Kent Yao	2b46662bd0	[SPARK-31111][SQL][TESTS] Fix interval output issue in ExtractBenchmark ### What changes were proposed in this pull request? fix the error caused by interval output in ExtractBenchmark ### Why are the changes needed? fix a bug in the test ```scala [info] Running case: cast to interval [error] Exception in thread "main" org.apache.spark.sql.AnalysisException: Cannot use interval type in the table schema.;; [error] OverwriteByExpression RelationV2[] noop-table, true, true [error] +- Project [(subtractdates(cast(cast(id#0L as timestamp) as date), -719162) + subtracttimestamps(cast(id#0L as timestamp), -30610249419876544)) AS ((CAST(CAST(id AS TIMESTAMP) AS DATE) - DATE '0001-01-01') + (CAST(id AS TIMESTAMP) - TIMESTAMP '1000-01-01 01:02:03.123456'))#2] [error] +- Range (1262304000, 1272304000, step=1, splits=Some(1)) [error] [error] at org.apache.spark.sql.catalyst.util.TypeUtils$.failWithIntervalType(TypeUtils.scala:106) [error] at org.apache.spark.sql.catalyst.analysis.CheckAnalysis.$anonfun$checkAnalysis$25(CheckAnalysis.scala:389) [error] at org.a ``` ### Does this PR introduce any user-facing change? no ### How was this patch tested? re-run benchmark Closes #27867 from yaooqinn/SPARK-31111. Authored-by: Kent Yao <yaooqinn@hotmail.com> Signed-off-by: Wenchen Fan <wenchen@databricks.com>	2020-03-11 20:13:59 +08:00
Liang-Chi Hsieh	15557a7d05	[SPARK-31071][SQL] Allow annotating non-null fields when encoding Java Beans ### What changes were proposed in this pull request? When encoding Java Beans to Spark DataFrame, respecting `javax.annotation.Nonnull` and producing non-null fields. ### Why are the changes needed? When encoding Java Beans to Spark DataFrame, non-primitive types are encoded as nullable fields. Although It works for most cases, it can be an issue under a few situations, e.g. the one described in the JIRA ticket when saving DataFrame to Avro format with non-null field. We should allow Spark users more flexibility when creating Spark DataFrame from Java Beans. Currently, Spark users cannot create DataFrame with non-nullable fields in the schema from beans with non-nullable properties. Although it is possible to project top-level columns with SQL expressions like `AssertNotNull` to make it non-null, for nested fields it is more tricky to do it similarly. ### Does this PR introduce any user-facing change? Yes. After this change, Spark users can use `javax.annotation.Nonnull` to annotate non-null fields in Java Beans when encoding beans to Spark DataFrame. ### How was this patch tested? Added unit test. Closes #27851 from viirya/SPARK-31071. Authored-by: Liang-Chi Hsieh <viirya@gmail.com> Signed-off-by: Wenchen Fan <wenchen@databricks.com>	2020-03-11 18:27:48 +08:00
Yuanjian Li	3493162c78	[SPARK-31030][SQL] Backward Compatibility for Parsing and formatting Datetime ### What changes were proposed in this pull request? In Spark version 2.4 and earlier, datetime parsing, formatting and conversion are performed by using the hybrid calendar (Julian + Gregorian). Since the Proleptic Gregorian calendar is de-facto calendar worldwide, as well as the chosen one in ANSI SQL standard, Spark 3.0 switches to it by using Java 8 API classes (the java.time packages that are based on ISO chronology ). The switching job is completed in SPARK-26651. But after the switching, there are some patterns not compatible between Java 8 and Java 7, Spark needs its own definition on the patterns rather than depends on Java API. In this PR, we achieve this by writing the document and shadow the incompatible letters. See more details in [SPARK-31030](https://issues.apache.org/jira/browse/SPARK-31030) ### Why are the changes needed? For backward compatibility. ### Does this PR introduce any user-facing change? No. After we define our own datetime parsing and formatting patterns, it's same to old Spark version. ### How was this patch tested? Existing and new added UT. Locally document test: ![image](https://user-images.githubusercontent.com/4833765/76064100-f6acc280-5fc3-11ea-9ef7-82e7dc074205.png) Closes #27830 from xuanyuanking/SPARK-31030. Authored-by: Yuanjian Li <xyliyuanjian@gmail.com> Signed-off-by: Wenchen Fan <wenchen@databricks.com>	2020-03-11 14:11:13 +08:00
Wenchen Fan	d5f5720efa	[SPARK-31070][SQL] make skew join split skewed partitions more evenly <!-- Thanks for sending a pull request! Here are some tips for you: 1. If this is your first time, please read our contributor guidelines: https://spark.apache.org/contributing.html 2. Ensure you have added or run the appropriate tests for your PR: https://spark.apache.org/developer-tools.html 3. If the PR is unfinished, add '[WIP]' in your PR title, e.g., '[WIP][SPARK-XXXX] Your PR title ...'. 4. Be sure to keep the PR description updated to reflect all changes. 5. Please write your PR title to summarize what this PR proposes. 6. If possible, provide a concise example to reproduce the issue for a faster review. 7. If you want to add a new configuration, please read the guideline first for naming configurations in 'core/src/main/scala/org/apache/spark/internal/config/ConfigEntry.scala'. --> ### What changes were proposed in this pull request? <!-- Please clarify what changes you are proposing. The purpose of this section is to outline the changes and how this PR fixes the issue. If possible, please consider writing useful notes for better and faster reviews in your PR. See the examples below. 1. If you refactor some codes with changing classes, showing the class hierarchy will help reviewers. 2. If you fix some SQL features, you can provide some references of other DBMSes. 3. If there is design documentation, please add the link. 4. If there is a discussion in the mailing list, please add the link. --> There are two problems when splitting skewed partitions: 1. It's impossible that we can't split the skewed partition, then we shouldn't create a skew join. 2. When splitting, it's possible that we create a partition for very small amount of data.. This PR fixes them 1. don't create `PartialReducerPartitionSpec` if we can't split. 2. merge small partitions to the previous partition. ### Why are the changes needed? <!-- Please clarify why the changes are needed. For instance, 1. If you propose a new API, clarify the use case for a new API. 2. If you fix a bug, you can clarify why it is a bug. --> make skew join split skewed partitions more evenly ### Does this PR introduce any user-facing change? <!-- If yes, please clarify the previous behavior and the change this PR proposes - provide the console output, description and/or an example to show the behavior difference if possible. If no, write 'No'. --> no ### How was this patch tested? <!-- If tests were added, say they were added here. Please make sure to add some test cases that check the changes thoroughly including negative and positive cases if possible. If it was tested in a way different from regular unit tests, please clarify how you tested step by step, ideally copy and paste-able, so that other reviewers can test and check, and descendants can verify in the future. If tests were not added, please describe why they were not added and/or why it was difficult to add. --> updated test Closes #27833 from cloud-fan/aqe. Authored-by: Wenchen Fan <wenchen@databricks.com> Signed-off-by: gatorsmile <gatorsmile@gmail.com>	2020-03-10 21:50:44 -07:00
Dongjoon Hyun	93def95b08	[SPARK-31095][BUILD] Upgrade netty-all to 4.1.47.Final ### What changes were proposed in this pull request? This PR aims to bring the bug fixes from the latest netty-all. ### Why are the changes needed? - 4.1.47.Final: https://github.com/netty/netty/milestone/222?closed=1 (15 patches or issues) - 4.1.46.Final: https://github.com/netty/netty/milestone/221?closed=1 (80 patches or issues) - 4.1.45.Final: https://github.com/netty/netty/milestone/220?closed=1 (23 patches or issues) - 4.1.44.Final: https://github.com/netty/netty/milestone/218?closed=1 (113 patches or issues) - 4.1.43.Final: https://github.com/netty/netty/milestone/217?closed=1 (63 patches or issues) ### Does this PR introduce any user-facing change? No. ### How was this patch tested? Pass the Jenkins with the existing tests. Closes #27869 from dongjoon-hyun/SPARK-31095. Authored-by: Dongjoon Hyun <dongjoon@apache.org> Signed-off-by: Dongjoon Hyun <dongjoon@apache.org>	2020-03-10 17:50:34 -07:00
Qianyang Yu	0f54dc7c03	[SPARK-30962][SQL][DOC] Documentation for Alter table command phase 2 ### What changes were proposed in this pull request? ### Why are the changes needed? Based on [JIRA 30962](https://issues.apache.org/jira/browse/SPARK-30962), we want to add all the support `Alter Table` syntax for V1 table. ### Does this PR introduce any user-facing change? Yes ### How was this patch tested? Before: The documentation looks like [Alter Table](https://github.com/apache/spark/pull/25590) After: <img width="850" alt="Screen Shot 2020-03-03 at 2 02 23 PM" src="https://user-images.githubusercontent.com/7550280/75824837-168c7e00-5d59-11ea-9751-d1dab0f5a892.png"> <img width="977" alt="Screen Shot 2020-03-03 at 2 02 41 PM" src="https://user-images.githubusercontent.com/7550280/75824859-21dfa980-5d59-11ea-8b49-3adf6eb55fc6.png"> <img width="1028" alt="Screen Shot 2020-03-03 at 2 02 59 PM" src="https://user-images.githubusercontent.com/7550280/75824884-2e640200-5d59-11ea-81ef-d77d0a8efee2.png"> <img width="864" alt="Screen Shot 2020-03-03 at 2 03 14 PM" src="https://user-images.githubusercontent.com/7550280/75824910-39b72d80-5d59-11ea-84d0-bffa2499f086.png"> <img width="823" alt="Screen Shot 2020-03-03 at 2 03 28 PM" src="https://user-images.githubusercontent.com/7550280/75824937-45a2ef80-5d59-11ea-932c-314924856834.png"> <img width="811" alt="Screen Shot 2020-03-03 at 2 03 42 PM" src="https://user-images.githubusercontent.com/7550280/75824965-4cc9fd80-5d59-11ea-815b-8c1ebad310b1.png"> <img width="827" alt="Screen Shot 2020-03-03 at 2 03 53 PM" src="https://user-images.githubusercontent.com/7550280/75824978-518eb180-5d59-11ea-8a55-2fa26376b9c1.png"> <img width="783" alt="Screen Shot 2020-03-03 at 2 04 03 PM" src="https://user-images.githubusercontent.com/7550280/75825001-5bb0b000-5d59-11ea-8dd9-dcfbfa1b4330.png"> Notes: Those syntaxes are not supported by v1 Table. - `ALTER TABLE .. RENAME COLUMN` - `ALTER TABLE ... DROP (COLUMN \| COLUMNS)` - `ALTER TABLE ... (ALTER \| CHANGE) COLUMN? alterColumnAction` only support change comments, not other actions: `datatype, position, (SET \| DROP) NOT NULL` - `ALTER TABLE .. CHANGE COLUMN?` - `ALTER TABLE .... REPLACE COLUMNS` - `ALTER TABLE ... RECOVER PARTITIONS` - Closes #27779 from kevinyu98/spark-30962-alterT. Authored-by: Qianyang Yu <qyu@us.ibm.com> Signed-off-by: Takeshi Yamamuro <yamamuro@apache.org>	2020-03-11 08:47:30 +09:00
yi.wu	34be83e08b	[SPARK-31037][SQL][FOLLOW-UP] Replace legacy ReduceNumShufflePartitions with CoalesceShufflePartitions in comment ### What changes were proposed in this pull request? Replace legacy `ReduceNumShufflePartitions` with `CoalesceShufflePartitions` in comment. ### Why are the changes needed? Rule `ReduceNumShufflePartitions` has renamed to `CoalesceShufflePartitions`, we should update related comment as well. ### Does this PR introduce any user-facing change? No. ### How was this patch tested? N/A. Closes #27865 from Ngone51/spark_31037_followup. Authored-by: yi.wu <yi.wu@databricks.com> Signed-off-by: Dongjoon Hyun <dongjoon@apache.org>	2020-03-10 11:09:36 -07:00
Kent Yao	3bd6ebff81	[SPARK-30189][SQL] Interval from year-month/date-time string should handle whitespaces ### What changes were proposed in this pull request? Currently, we parse interval from multi units strings or from date-time/year-month pattern strings, the former handles all whitespace, the latter not or even spaces. ### Why are the changes needed? behavior consistency ### Does this PR introduce any user-facing change? yes, interval in date-time/year-month like ``` select interval '\n-\t10\t 12:34:46.789\t' day to second -- !query 126 schema struct<INTERVAL '-10 days -12 hours -34 minutes -46.789 seconds':interval> -- !query 126 output -10 days -12 hours -34 minutes -46.789 seconds ``` is valid now. ### How was this patch tested? add ut. Closes #26815 from yaooqinn/SPARK-30189. Authored-by: Kent Yao <yaooqinn@hotmail.com> Signed-off-by: Wenchen Fan <wenchen@databricks.com>	2020-03-10 22:08:58 +08:00
Terry Kim	294f6056fa	[SPARK-31078][SQL] Respect aliases in output ordering ### What changes were proposed in this pull request? Currently, in the following scenario, an unnecessary `Sort` node is introduced: ```scala withSQLConf(SQLConf.AUTO_BROADCASTJOIN_THRESHOLD.key -> "0") { val df = (0 until 20).toDF("i").as("df") df.repartition(8, df("i")).write.format("parquet") .bucketBy(8, "i").sortBy("i").saveAsTable("t") val t1 = spark.table("t") val t2 = t1.selectExpr("i as ii") t1.join(t2, t1("i") === t2("ii")).explain } ``` ``` == Physical Plan == (3) SortMergeJoin [i#8], [ii#10], Inner :- (1) Project [i#8] : +- (1) Filter isnotnull(i#8) : +- (1) ColumnarToRow : +- FileScan parquet default.t[i#8] Batched: true, DataFilters: [isnotnull(i#8)], Format: Parquet, Location: InMemoryFileIndex[file:/..., PartitionFilters: [], PushedFilters: [IsNotNull(i)], ReadSchema: struct<i:int>, SelectedBucketsCount: 8 out of 8 +- (2) Sort [ii#10 ASC NULLS FIRST], false, 0 <==== UNNECESSARY +- (2) Project [i#8 AS ii#10] +- (2) Filter isnotnull(i#8) +- (2) ColumnarToRow +- FileScan parquet default.t[i#8] Batched: true, DataFilters: [isnotnull(i#8)], Format: Parquet, Location: InMemoryFileIndex[file:/..., PartitionFilters: [], PushedFilters: [IsNotNull(i)], ReadSchema: struct<i:int>, SelectedBucketsCount: 8 out of 8 ``` Notice that `Sort [ii#10 ASC NULLS FIRST], false, 0` is introduced even though the underlying data is already sorted. This is because `outputOrdering` doesn't handle aliases correctly. This PR proposes to fix this issue. ### Why are the changes needed? To better handle aliases in `outputOrdering`. ### Does this PR introduce any user-facing change? Yes, now with the fix, the `explain` prints out the following: ``` == Physical Plan == (3) SortMergeJoin [i#8], [ii#10], Inner :- (1) Project [i#8] : +- (1) Filter isnotnull(i#8) : +- (1) ColumnarToRow : +- FileScan parquet default.t[i#8] Batched: true, DataFilters: [isnotnull(i#8)], Format: Parquet, Location: InMemoryFileIndex[file:/..., PartitionFilters: [], PushedFilters: [IsNotNull(i)], ReadSchema: struct<i:int>, SelectedBucketsCount: 8 out of 8 +- (2) Project [i#8 AS ii#10] +- (2) Filter isnotnull(i#8) +- *(2) ColumnarToRow +- FileScan parquet default.t[i#8] Batched: true, DataFilters: [isnotnull(i#8)], Format: Parquet, Location: InMemoryFileIndex[file:/..., PartitionFilters: [], PushedFilters: [IsNotNull(i)], ReadSchema: struct<i:int>, SelectedBucketsCount: 8 out of 8 ``` ### How was this patch tested? Tests added. Closes #27842 from imback82/alias_aware_sort_order. Authored-by: Terry Kim <yuminkim@gmail.com> Signed-off-by: Wenchen Fan <wenchen@databricks.com>	2020-03-10 20:15:48 +08:00
Eric Wu	15df2a3f40	[SPARK-31079][SQL] Logging QueryExecutionMetering in RuleExecutor logger ### What changes were proposed in this pull request? RuleExecutor already support metering for analyzer/optimizer rules. By providing such information in `PlanChangeLogger`, user can get more information when debugging rule changes . This PR enhanced `PlanChangeLogger` to display RuleExecutor metrics. This can be easily done by calling the existing API `resetMetrics` and `dumpTimeSpent`, but there might be conflicts if user is also collecting total metrics of a sql job. Thus I introduced `QueryExecutionMetrics`, as the snapshot of `QueryExecutionMetering`, to better support this feature. Information added to `PlanChangeLogger` ``` === Metrics of Executed Rules === Total number of runs: 554 Total time: 0.107756568 seconds Total number of effective runs: 11 Total time of effective runs: 0.047615486 seconds ``` ### Why are the changes needed? Provide better plan change debugging user experience ### Does this PR introduce any user-facing change? Only add more debugging info of `planChangeLog`, default log level is TRACE. ### How was this patch tested? Update existing tests to verify the new logs Closes #27846 from Eric5553/ExplainRuleExecMetrics. Authored-by: Eric Wu <492960551@qq.com> Signed-off-by: Wenchen Fan <wenchen@databricks.com>	2020-03-10 19:08:59 +08:00
beliefer	8ee41f3576	[SPARK-30992][DSTREAMS] Arrange scattered config of streaming module ### What changes were proposed in this pull request? I found a lot scattered config in `Streaming`.I think should arrange these config in unified position. ### Why are the changes needed? Arrange scattered config ### Does this PR introduce any user-facing change? No ### How was this patch tested? Exists UT Closes #27744 from beliefer/arrange-scattered-streaming-config. Authored-by: beliefer <beliefer@163.com> Signed-off-by: HyukjinKwon <gurwls223@apache.org>	2020-03-10 18:04:09 +09:00
HyukjinKwon	815c7929c2	[SPARK-31065][SQL] Match schema_of_json to the schema inference of JSON data source ### What changes were proposed in this pull request? This PR proposes two things: 1. Convert `null` to `string` type during schema inference of `schema_of_json` as JSON datasource does. This is a bug fix as well because `null` string is not the proper DDL formatted string and it is unable for SQL parser to recognise it as a type string. We should match it to JSON datasource and return a string type so `schema_of_json` returns a proper DDL formatted string. 2. Let `schema_of_json` respect `dropFieldIfAllNull` option during schema inference. ### Why are the changes needed? To let `schema_of_json` return a proper DDL formatted string, and respect `dropFieldIfAllNull` option. ### Does this PR introduce any user-facing change? Yes, it does. ```scala import collection.JavaConverters._ import org.apache.spark.sql.functions._ spark.range(1).select(schema_of_json(lit("""{"id": ""}"""))).show() spark.range(1).select(schema_of_json(lit("""{"id": "a", "drop": {"drop": null}}"""), Map("dropFieldIfAllNull" -> "true").asJava)).show(false) ``` Before: ``` struct<id:null> struct<drop:struct<drop:null>,id:string> ``` After: ``` struct<id:string> struct<id:string> ``` ### How was this patch tested? Manually tested, and unittests were added. Closes #27854 from HyukjinKwon/SPARK-31065. Authored-by: HyukjinKwon <gurwls223@apache.org> Signed-off-by: Dongjoon Hyun <dongjoon@apache.org>	2020-03-10 00:33:32 -07:00
maryannxue	de6d9e4307	[SPARK-31096][SQL] Replace `Array` with `Seq` in AQE `CustomShuffleReaderExec` ### What changes were proposed in this pull request? This PR changes the type of `CustomShuffleReaderExec`'s `partitionSpecs` from `Array` to `Seq`, since `Array` compares references not values for equality, which could lead to potential plan reuse problem. ### Why are the changes needed? Unlike `Seq`, `Array` compares references not values for equality, which could lead to potential plan reuse problem. ### Does this PR introduce any user-facing change? No. ### How was this patch tested? Passes existing UTs. Closes #27857 from maryannxue/aqe-customreader-fix. Authored-by: maryannxue <maryannxue@apache.org> Signed-off-by: Wenchen Fan <wenchen@databricks.com>	2020-03-10 14:15:44 +08:00
Yuchen Huo	a22994333a	[SPARK-30902][SQL][FOLLOW-UP] Allow ReplaceTableAsStatement to have none provider ### What changes were proposed in this pull request? This is a follow up for https://github.com/apache/spark/pull/27650 where allow None provider for create table. Here we are doing the same thing for ReplaceTable. ### Why are the changes needed? Although currently the ASTBuilder doesn't seem to allow `replace` without `USING` clause. This would allow `DataFrameWriterV2` to use the statements instead of commands directly. ### Does this PR introduce any user-facing change? No. ### How was this patch tested? Existing tests Closes #27838 from yuchenhuo/SPARK-30902. Authored-by: Yuchen Huo <yuchen.huo@databricks.com> Signed-off-by: Wenchen Fan <wenchen@databricks.com>	2020-03-10 11:37:31 +08:00
Thomas Graves	e807118eef	[SPARK-31055][DOCS] Update config docs for shuffle local host reads to have dep on external shuffle service ### What changes were proposed in this pull request? with SPARK-27651 we now support host local reads for shuffle, but only when external shuffle service is enabled. Update the config docs to state that. ### Why are the changes needed? clarify dependency ### Does this PR introduce any user-facing change? no ### How was this patch tested? n/a Closes #27812 from tgravescs/SPARK-27651-follow. Authored-by: Thomas Graves <tgraves@nvidia.com> Signed-off-by: Dongjoon Hyun <dongjoon@apache.org>	2020-03-09 12:17:59 -07:00
Liang-Chi Hsieh	d21aab403a	[SPARK-30941][PYSPARK] Add a note to asDict to document its behavior when there are duplicate fields ### What changes were proposed in this pull request? Adding a note to document `Row.asDict` behavior when there are duplicate fields. ### Why are the changes needed? When a row contains duplicate fields, `asDict` and `_get_item_` behaves differently. We should document it to let users know the difference explicitly. ### Does this PR introduce any user-facing change? No. Only document change. ### How was this patch tested? Existing test. Closes #27853 from viirya/SPARK-30941. Authored-by: Liang-Chi Hsieh <viirya@gmail.com> Signed-off-by: Dongjoon Hyun <dongjoon@apache.org>	2020-03-09 11:06:45 -07:00
Huaxin Gao	b6b0343e3e	[SPARK-30929][ML] ML, GraphX 3.0 QA: API: New Scala APIs, docs ### What changes were proposed in this pull request? Auditing new ML Scala APIs introduced in 3.0. Fix found issues. ### Why are the changes needed? ### Does this PR introduce any user-facing change? Yes. Some doc changes ### How was this patch tested? Existing tests Closes #27818 from huaxingao/spark-30929. Authored-by: Huaxin Gao <huaxing@us.ibm.com> Signed-off-by: Sean Owen <srowen@gmail.com>	2020-03-09 09:11:21 -05:00
yi.wu	ef51ff9dc8	[SPARK-31082][CORE] MapOutputTrackerMaster.getMapLocation should handle last mapIndex correctly ### What changes were proposed in this pull request? In `getMapLocation`, change the condition from `...endMapIndex < statuses.length` to `...endMapIndex <= statuses.length`. ### Why are the changes needed? `endMapIndex` is exclusive, we should include it when comparing to `statuses.length`. Otherwise, we can't get the location for last mapIndex. ### Does this PR introduce any user-facing change? No. ### How was this patch tested? Updated existed test. Closes #27850 from Ngone51/fix_getmaploction. Authored-by: yi.wu <yi.wu@databricks.com> Signed-off-by: Wenchen Fan <wenchen@databricks.com>	2020-03-09 15:53:34 +08:00
Kousuke Saruta	068bdd4415	[SPARK-31073][WEBUI] Add "shuffle write time" to task metrics summary in StagePage ### What changes were proposed in this pull request? I've applied following changed to StagePage. 1. Added `Shuffle Write Time` to task metrics summary. 2. Added checkbox for `Shuffle Write Time` as an additional metrics. 3. Renamed `Write Time` column in task table to `Shuffle Write Time` and let it as an additional column. ### Why are the changes needed? Task metrics summary doesn't show `Shuffle Write Time` even though it shows `Shuffle Read Blocked Time`. `Shuffle Read Blocked Time` is let as an additional metrics so I also let `Shuffle Write Time` as an other additional metrics. ### Does this PR introduce any user-facing change? Yes. After this change, task metrics summary can show `Shuffle Write Time` and its visibility is controlled by a checkbox. ![additional-metrics-after](https://user-images.githubusercontent.com/4736016/76101844-677acb80-6012-11ea-9923-d95d852c775b.png) ![task-summary-after](https://user-images.githubusercontent.com/4736016/76101856-6ea1d980-6012-11ea-9670-3cf0ecd6faff.png) `Write Time` column is already shown in task table but the title is ambiguous so I've renamed it as `Shuffle Write Time`. After this change, this column is also additional column like `Shuffle Read Blocked Time`. ![tasks-table-after](https://user-images.githubusercontent.com/4736016/76102216-00a9e200-6013-11ea-9d51-1a6ce2abb0b9.png) ### How was this patch tested? I've tested manually using following code and confirm the UI. `sc.parallelize(1 to 1000).map(x => (x,x)).reduceByKey(_+_).collect` Closes #27837 from sarutak/write-time. Authored-by: Kousuke Saruta <sarutak@oss.nttdata.com> Signed-off-by: Dongjoon Hyun <dongjoon@apache.org>	2020-03-08 20:20:39 -07:00

... 14 15 16 17 18 ...

27485 commits