ODIn/spark-instrumented-optimizer

Author	SHA1	Message	Date
Huaxin Gao	1a7f9649b6	[SPARK-31305][SQL][DOCS] Add a page to list all commands in SQL Reference ### What changes were proposed in this pull request? Add a page to list all commands in SQL Reference... ### Why are the changes needed? so it's easier for user to find a specific command. ### Does this PR introduce any user-facing change? before: ![image](https://user-images.githubusercontent.com/13592258/77938658-ec03e700-726a-11ea-983c-7a559cc0aae2.png) after: ![image](https://user-images.githubusercontent.com/13592258/77937899-d3df9800-7269-11ea-85db-749a9521576a.png) ![image](https://user-images.githubusercontent.com/13592258/77937924-db9f3c80-7269-11ea-9441-7603feee421c.png) Also move ```use database``` from query category to ddl category. ### How was this patch tested? Manually build and check Closes #28074 from huaxingao/list-all. Authored-by: Huaxin Gao <huaxing@us.ibm.com> Signed-off-by: Takeshi Yamamuro <yamamuro@apache.org>	2020-04-01 08:42:15 +09:00
Qianyang Yu	e65c21e093	[SPARK-31304][ML][EXAMPLES] Add examples for ml.stat.ANOVATest ### What changes were proposed in this pull request? Add ANOVATest example for ml.stat.ANOVATest in python/java/scala ### Why are the changes needed? Improve ML example ### Does this PR introduce any user-facing change? No ### How was this patch tested? manually run the example Closes #28073 from kevinyu98/add-ANOVA-example. Authored-by: Qianyang Yu <qyu@us.ibm.com> Signed-off-by: Sean Owen <srowen@gmail.com>	2020-03-31 16:33:26 -05:00
Wenchen Fan	34c7ec8e0c	[SPARK-31253][SQL] Add metrics to AQE shuffle reader <!-- Thanks for sending a pull request! Here are some tips for you: 1. If this is your first time, please read our contributor guidelines: https://spark.apache.org/contributing.html 2. Ensure you have added or run the appropriate tests for your PR: https://spark.apache.org/developer-tools.html 3. If the PR is unfinished, add '[WIP]' in your PR title, e.g., '[WIP][SPARK-XXXX] Your PR title ...'. 4. Be sure to keep the PR description updated to reflect all changes. 5. Please write your PR title to summarize what this PR proposes. 6. If possible, provide a concise example to reproduce the issue for a faster review. 7. If you want to add a new configuration, please read the guideline first for naming configurations in 'core/src/main/scala/org/apache/spark/internal/config/ConfigEntry.scala'. --> ### What changes were proposed in this pull request? <!-- Please clarify what changes you are proposing. The purpose of this section is to outline the changes and how this PR fixes the issue. If possible, please consider writing useful notes for better and faster reviews in your PR. See the examples below. 1. If you refactor some codes with changing classes, showing the class hierarchy will help reviewers. 2. If you fix some SQL features, you can provide some references of other DBMSes. 3. If there is design documentation, please add the link. 4. If there is a discussion in the mailing list, please add the link. --> Add SQL metrics to the AQE shuffle reader (`CustomShuffleReaderExec`) ### Why are the changes needed? <!-- Please clarify why the changes are needed. For instance, 1. If you propose a new API, clarify the use case for a new API. 2. If you fix a bug, you can clarify why it is a bug. --> to be more UI friendly ### Does this PR introduce any user-facing change? <!-- If yes, please clarify the previous behavior and the change this PR proposes - provide the console output, description and/or an example to show the behavior difference if possible. If no, write 'No'. --> No ### How was this patch tested? <!-- If tests were added, say they were added here. Please make sure to add some test cases that check the changes thoroughly including negative and positive cases if possible. If it was tested in a way different from regular unit tests, please clarify how you tested step by step, ideally copy and paste-able, so that other reviewers can test and check, and descendants can verify in the future. If tests were not added, please describe why they were not added and/or why it was difficult to add. --> new test Closes #28022 from cloud-fan/metrics. Authored-by: Wenchen Fan <wenchen@databricks.com> Signed-off-by: gatorsmile <gatorsmile@gmail.com>	2020-03-31 13:03:52 -07:00
yi.wu	590b9a0132	[SPARK-31010][SQL][FOLLOW-UP] Add Java UDF suggestion in error message of untyped Scala UDF ### What changes were proposed in this pull request? Added Java UDF suggestion in the in error message of untyped Scala UDF. ### Why are the changes needed? To help user migrate their use case from deprecate untyped Scala UDF to other supported UDF. ### Does this PR introduce any user-facing change? No. It haven't been released. ### How was this patch tested? Pass Jenkins. Closes #28070 from Ngone51/spark_31010. Authored-by: yi.wu <yi.wu@databricks.com> Signed-off-by: Wenchen Fan <wenchen@databricks.com>	2020-03-31 17:35:26 +00:00
Jungtaek Lim (HeartSaVioR)	2a6aa8e87b	[SPARK-31312][SQL] Cache Class instance for the UDF instance in HiveFunctionWrapper ### What changes were proposed in this pull request? This patch proposes to cache Class instance for the UDF instance in HiveFunctionWrapper to fix the case where Hive simple UDF is somehow transformed (expression is copied) and evaluated later with another classloader (for the case current thread context classloader is somehow changed). In this case, Spark throws CNFE as of now. It's only occurred for Hive simple UDF, as HiveFunctionWrapper caches the UDF instance whereas it doesn't do for `UDF` type. The comment says Spark has to create instance every time for UDF, so we cannot simply do the same. This patch caches Class instance instead, and switch current thread context classloader to which loads the Class instance. This patch extends the test boundary as well. We only tested with GenericUDTF for SPARK-26560, and this patch actually requires only UDF. But to avoid regression for other types as well, this patch adds all available types (UDF, GenericUDF, AbstractGenericUDAFResolver, UDAF, GenericUDTF) into the boundary of tests. Credit to cloud-fan as he discovered the problem and proposed the solution. ### Why are the changes needed? Above section describes why it's a bug and how it's fixed. ### Does this PR introduce any user-facing change? No. ### How was this patch tested? New UTs added. Closes #28079 from HeartSaVioR/SPARK-31312. Authored-by: Jungtaek Lim (HeartSaVioR) <kabhwan.opensource@gmail.com> Signed-off-by: Wenchen Fan <wenchen@databricks.com>	2020-03-31 16:17:26 +00:00
Wenchen Fan	8b01473e8b	[SPARK-31230][SQL] Use statement plans in DataFrameWriter(V2) ### What changes were proposed in this pull request? Create statement plans in `DataFrameWriter(V2)`, like the SQL API. ### Why are the changes needed? It's better to leave all the resolution work to the analyzer. ### Does this PR introduce any user-facing change? no ### How was this patch tested? existing tests Closes #27992 from cloud-fan/statement. Authored-by: Wenchen Fan <wenchen@databricks.com> Signed-off-by: Wenchen Fan <wenchen@databricks.com>	2020-03-31 23:19:46 +08:00
Yuanjian Li	07c50784d3	[SPARK-31314][CORE] Revert SPARK-29285 to fix shuffle regression caused by creating temporary file eagerly ### What changes were proposed in this pull request? This reverts commit `8cf76f8d61`. #25962 ### Why are the changes needed? In SPARK-29285, we change to create shuffle temporary eagerly. This is helpful for not to fail the entire task in the scenario of occasional disk failure. But for the applications that many tasks don't actually create shuffle files, it caused overhead. See the below benchmark: Env: Spark local-cluster[2, 4, 19968], each queries run 5 round, each round 5 times. Data: TPC-DS scale=99 generate by spark-tpcds-datagen Results: \| \| Base \| Revert \| \|-----\|---------------------------------------------------------------------------------------------\|---------------------------------------------------------------------------------------------\| \| Q20 \| Vector(4.096865667, 2.76231748, 2.722007606, 2.514433591, 2.400373579) Median 2.722007606 \| Vector(3.763185446, 2.586498463, 2.593472842, 2.320522846, 2.224627274) Median 2.586498463 \| \| Q33 \| Vector(5.872176321, 4.854397586, 4.568787136, 4.393378146, 4.423996818) Median 4.568787136 \| Vector(5.38746785, 4.361236877, 4.082311276, 3.867206824, 3.783188024) Median 4.082311276 \| \| Q52 \| Vector(3.978870321, 3.225437871, 3.282411608, 2.869674887, 2.644490664) Median 3.225437871 \| Vector(4.000381522, 3.196025108, 3.248787619, 2.767444508, 2.606163423) Median 3.196025108 \| \| Q56 \| Vector(6.238045133, 4.820535173, 4.609965579, 4.313509894, 4.221256227) Median 4.609965579 \| Vector(6.241611339, 4.225592467, 4.195202502, 3.757085755, 3.657525982) Median 4.195202502 \| ### Does this PR introduce any user-facing change? No ### How was this patch tested? Existing tests. Closes #28072 from xuanyuanking/SPARK-29285-revert. Authored-by: Yuanjian Li <xyliyuanjian@gmail.com> Signed-off-by: Wenchen Fan <wenchen@databricks.com>	2020-03-31 19:01:08 +08:00
Maxim Gekk	bb0b416f0b	[SPARK-31297][SQL] Speed up dates rebasing ### What changes were proposed in this pull request? In the PR, I propose to replace current implementation of the `rebaseGregorianToJulianDays()` and `rebaseJulianToGregorianDays()` functions in `DateTimeUtils` by new one which is based on the fact that difference between Proleptic Gregorian and the hybrid (Julian+Gregorian) calendars was changed only 14 times for entire supported range of valid dates `[0001-01-01, 9999-12-31]`: \| date \| Proleptic Greg. days \| Hybrid (Julian+Greg) days \| diff\| \| ---- \| ----\|----\|----\| \|0001-01-01\|-719162\|-719164\|-2\| \|0100-03-01\|-682944\|-682945\|-1\| \|0200-03-01\|-646420\|-646420\|0\| \|0300-03-01\|-609896\|-609895\|1\| \|0500-03-01\|-536847\|-536845\|2\| \|0600-03-01\|-500323\|-500320\|3\| \|0700-03-01\|-463799\|-463795\|4\| \|0900-03-01\|-390750\|-390745\|5\| \|1000-03-01\|-354226\|-354220\|6\| \|1100-03-01\|-317702\|-317695\|7\| \|1300-03-01\|-244653\|-244645\|8\| \|1400-03-01\|-208129\|-208120\|9\| \|1500-03-01\|-171605\|-171595\|10\| \|1582-10-15\|-141427\|-141427\|0\| For the given days since the epoch, the proposed implementation finds the range of days which the input days belongs to, and adds the diff in days between calendars to the input. The result is rebased days since the epoch in the target calendar. For example, if need to rebase -650000 days from Proleptic Gregorian calendar to the hybrid calendar. In that case, the input falls to the bucket [-682944, -646420), the diff associated with the range is -1. To get the rebased days in Julian calendar, we should add -1 to -650000, and the result is -650001. ### Why are the changes needed? To make dates rebasing faster. ### Does this PR introduce any user-facing change? No, the results should be the same for valid range of the `DATE` type `[0001-01-01, 9999-12-31]`. ### How was this patch tested? - Added 2 tests to `DateTimeUtilsSuite` for the `rebaseGregorianToJulianDays()` and `rebaseJulianToGregorianDays()` functions. The tests check that results of old and new implementation (optimized version) are the same for all supported dates. - Re-run `DateTimeRebaseBenchmark` on: \| Item \| Description \| \| ---- \| ----\| \| Region \| us-west-2 (Oregon) \| \| Instance \| r3.xlarge \| \| AMI \| ubuntu/images/hvm-ssd/ubuntu-bionic-18.04-amd64-server-20190722.1 (ami-06f2f779464715dc5) \| \| Java \| OpenJDK8/11 \| Closes #28067 from MaxGekk/optimize-rebasing. Lead-authored-by: Maxim Gekk <max.gekk@gmail.com> Co-authored-by: Max Gekk <max.gekk@gmail.com> Signed-off-by: Wenchen Fan <wenchen@databricks.com>	2020-03-31 17:38:47 +08:00
HyukjinKwon	4d4c3e76f6	Revert "[SPARK-30879][DOCS] Refine workflow for building docs" This reverts commit `7892f88f84`.	2020-03-31 16:11:59 +09:00
Ben Ryves	fa37856710	[SPARK-31306][DOCS] update rand() function documentation to indicate exclusive upper bound ### What changes were proposed in this pull request? A small documentation change to clarify that the `rand()` function produces values in `[0.0, 1.0)`. ### Why are the changes needed? `rand()` uses `Rand()` - which generates values in [0, 1) ([documented here](`a1dbcd13a3/sql/catalyst/src/main/scala/org/apache/spark/sql/catalyst/expressions/randomExpressions.scala (L71)`)). The existing documentation suggests that 1.0 is a possible value returned by rand (i.e for a distribution written as `X ~ U(a, b)`, x can be a or b, so `U[0.0, 1.0]` suggests the value returned could include 1.0). ### Does this PR introduce any user-facing change? Only documentation changes. ### How was this patch tested? Documentation changes only. Closes #28071 from Smeb/master. Authored-by: Ben Ryves <benjamin.ryves@getyourguide.com> Signed-off-by: HyukjinKwon <gurwls223@apache.org>	2020-03-31 15:16:17 +09:00
beliefer	47c810f8ae	[SPARK-31279][SQL][DOC] Add version information to the configuration of Hive ### What changes were proposed in this pull request? Add version information to the configuration of `Hive`. I sorted out some information show below. Item name \| Since version \| JIRA ID \| Commit ID \| Note -- \| -- \| -- \| -- \| -- spark.sql.hive.metastore.version \| 1.4.0 \| SPARK-6908 \| 05454fd8aef75b129cbbd0288f5089c5259f4a15#diff-ff50aea397a607b79df9bec6f2a841db \| spark.sql.hive.version \| 1.1.1 \| SPARK-3971 \| 64945f868443fbc59cb34b34c16d782dda0fb63d#diff-12fa2178364a810b3262b30d8d48aa2d \| spark.sql.hive.metastore.jars \| 1.4.0 \| SPARK-6908 \| 05454fd8aef75b129cbbd0288f5089c5259f4a15#diff-ff50aea397a607b79df9bec6f2a841db \| spark.sql.hive.convertMetastoreParquet \| 1.1.1 \| SPARK-2406 \| cc4015d2fa3785b92e6ab079b3abcf17627f7c56#diff-ff50aea397a607b79df9bec6f2a841db \| spark.sql.hive.convertMetastoreParquet.mergeSchema \| 1.3.1 \| SPARK-6575 \| 778c87686af0c04df9dfe144b8f744f271a988ad#diff-ff50aea397a607b79df9bec6f2a841db \| spark.sql.hive.convertMetastoreOrc \| 2.0.0 \| SPARK-14070 \| 1e886159849e3918445d3fdc3c4cef86c6c1a236#diff-ff50aea397a607b79df9bec6f2a841db \| spark.sql.hive.convertInsertingPartitionedTable \| 3.0.0 \| SPARK-28573 \| d5688dc732890923c326f272b0c18c329a69459a#diff-842e3447fc453de26c706db1cac8f2c4 \| spark.sql.hive.convertMetastoreCtas \| 3.0.0 \| SPARK-25271 \| 5ad03607d1487e7ab3e3b6d00eef9c4028ed4975#diff-842e3447fc453de26c706db1cac8f2c4 \| spark.sql.hive.metastore.sharedPrefixes \| 1.4.0 \| SPARK-7491 \| a8556086d33cb993fab0ae2751e31455e6c664ab#diff-ff50aea397a607b79df9bec6f2a841db \| spark.sql.hive.metastore.barrierPrefixes \| 1.4.0 \| SPARK-7491 \| a8556086d33cb993fab0ae2751e31455e6c664ab#diff-ff50aea397a607b79df9bec6f2a841db \| spark.sql.hive.thriftServer.async \| 1.5.0 \| SPARK-6964 \| eb19d3f75cbd002f7e72ce02017a8de67f562792#diff-ff50aea397a607b79df9bec6f2a841db \| ### Why are the changes needed? Supplemental configuration version information. ### Does this PR introduce any user-facing change? 'No'. ### How was this patch tested? Exists UT Closes #28042 from beliefer/add-version-to-hive-config. Authored-by: beliefer <beliefer@163.com> Signed-off-by: HyukjinKwon <gurwls223@apache.org>	2020-03-31 12:35:01 +09:00
beliefer	4fc8ee74fc	[SPARK-31295][DOC] Supplement version for configuration appear in doc ### What changes were proposed in this pull request? This PR supplements version for configuration appear in docs. I sorted out some information show below. docs/spark-standalone.md Item name \| Since version \| JIRA ID \| Commit ID \| Note -- \| -- \| -- \| -- \| -- spark.deploy.retainedApplications \| 0.8.0 \| None \| 46eecd110a4017ea0c86cbb1010d0ccd6a5eb2ef#diff-29dffdccd5a7f4c8b496c293e87c8668 \| spark.deploy.retainedDrivers \| 1.1.0 \| None \| 7446f5ff93142d2dd5c79c63fa947f47a1d4db8b#diff-29dffdccd5a7f4c8b496c293e87c8668 \| spark.deploy.spreadOut \| 0.6.1 \| None \| bb2b9ff37cd2503cc6ea82c5dd395187b0910af0#diff-0e7ae91819fc8f7b47b0f97be7116325 \| spark.deploy.defaultCores \| 0.9.0 \| None \| d8bcc8e9a095c1b20dd7a17b6535800d39bff80e#diff-29dffdccd5a7f4c8b496c293e87c8668 \| spark.deploy.maxExecutorRetries \| 1.6.3 \| SPARK-16956 \| ace458f0330f22463ecf7cbee7c0465e10fba8a8#diff-29dffdccd5a7f4c8b496c293e87c8668 \| spark.worker.resource.{resourceName}.amount \| 3.0.0 \| SPARK-27371 \| cbad616d4cb0c58993a88df14b5e30778c7f7e85#diff-d25032e4a3ae1b85a59e4ca9ccf189a8 \| spark.worker.resource.{resourceName}.discoveryScript \| 3.0.0 \| SPARK-27371 \| cbad616d4cb0c58993a88df14b5e30778c7f7e85#diff-d25032e4a3ae1b85a59e4ca9ccf189a8 \| spark.worker.resourcesFile \| 3.0.0 \| SPARK-27369 \| 7cbe01e8efc3f6cd3a0cac4bcfadea8fcc74a955#diff-b2fc8d6ab7ac5735085e2d6cfacb95da \| spark.shuffle.service.db.enabled \| 3.0.0 \| SPARK-26288 \| 8b0aa59218c209d39cbba5959302d8668b885cf6#diff-6bdad48cfc34314e89599655442ff210 \| spark.storage.cleanupFilesAfterExecutorExit \| 2.4.0 \| SPARK-24340 \| 8ef167a5f9ba8a79bb7ca98a9844fe9cfcfea060#diff-916ca56b663f178f302c265b7ef38499 \| spark.deploy.recoveryMode \| 0.8.1 \| None \| d66c01f2b6defb3db6c1be99523b734a4d960532#diff-29dffdccd5a7f4c8b496c293e87c8668 \| spark.deploy.recoveryDirectory \| 0.8.1 \| None \| d66c01f2b6defb3db6c1be99523b734a4d960532#diff-29dffdccd5a7f4c8b496c293e87c8668 \| docs/sql-data-sources-avro.md Item name \| Since version \| JIRA ID \| Commit ID \| Note -- \| -- \| -- \| -- \| -- spark.sql.legacy.replaceDatabricksSparkAvro.enabled \| 2.4.0 \| SPARK-25129 \| ac0174e55af2e935d41545721e9f430c942b3a0c#diff-9a6b543db706f1a90f790783d6930a13 \| spark.sql.avro.compression.codec \| 2.4.0 \| SPARK-24881 \| 0a0f68bae6c0a1bf30184b1e9ac6bf3805bd7511#diff-9a6b543db706f1a90f790783d6930a13 \| spark.sql.avro.deflate.level \| 2.4.0 \| SPARK-24881 \| 0a0f68bae6c0a1bf30184b1e9ac6bf3805bd7511#diff-9a6b543db706f1a90f790783d6930a13 \| docs/sql-data-sources-orc.md Item name \| Since version \| JIRA ID \| Commit ID \| Note -- \| -- \| -- \| -- \| -- spark.sql.orc.impl \| 2.3.0 \| SPARK-20728 \| 326f1d6728a7734c228d8bfaa69442a1c7b92e9b#diff-9a6b543db706f1a90f790783d6930a13 \| spark.sql.orc.enableVectorizedReader \| 2.3.0 \| SPARK-16060 \| 60f6b994505e3f82091a04eed2dc0a9e8bd523ce#diff-9a6b543db706f1a90f790783d6930a13 \| docs/sql-data-sources-parquet.md Item name \| Since version \| JIRA ID \| Commit ID \| Note -- \| -- \| -- \| -- \| -- spark.sql.parquet.binaryAsString \| 1.1.1 \| SPARK-2927 \| de501e169f24e4573747aec85b7651c98633c028#diff-41ef65b9ef5b518f77e2a03559893f4d \| spark.sql.parquet.int96AsTimestamp \| 1.3.0 \| SPARK-4987 \| 67d52207b5cf2df37ca70daff2a160117510f55e#diff-41ef65b9ef5b518f77e2a03559893f4d \| spark.sql.parquet.compression.codec \| 1.1.1 \| SPARK-3131 \| 3a9d874d7a46ab8b015631d91ba479d9a0ba827f#diff-41ef65b9ef5b518f77e2a03559893f4d \| spark.sql.parquet.filterPushdown \| 1.2.0 \| SPARK-4391 \| 576688aa2a19bd4ba239a2b93af7947f983e5124#diff-41ef65b9ef5b518f77e2a03559893f4d \| spark.sql.hive.convertMetastoreParquet \| 1.1.1 \| SPARK-2406 \| cc4015d2fa3785b92e6ab079b3abcf17627f7c56#diff-ff50aea397a607b79df9bec6f2a841db \| spark.sql.parquet.mergeSchema \| 1.5.0 \| SPARK-8690 \| 246265f2bb056d5e9011d3331b809471a24ff8d7#diff-41ef65b9ef5b518f77e2a03559893f4d \| spark.sql.parquet.writeLegacyFormat \| 1.6.0 \| SPARK-10400 \| 01cd688f5245cbb752863100b399b525b31c3510#diff-41ef65b9ef5b518f77e2a03559893f4d \| ### Why are the changes needed? Supplemental configuration version information. ### Does this PR introduce any user-facing change? 'No'. ### How was this patch tested? Jenkins test Closes #28064 from beliefer/supplement-doc-for-data-sources. Authored-by: beliefer <beliefer@163.com> Signed-off-by: HyukjinKwon <gurwls223@apache.org>	2020-03-31 12:33:46 +09:00
beliefer	fc5d67fe22	[SPARK-31282][DOC] Supplement version for configuration appear in security doc ### What changes were proposed in this pull request? This PR supplements version for configuration appear in security doc. I sorted out some information show below. Item name \| Since version \| JIRA ID \| Commit ID \| Note -- \| -- \| -- \| -- \| -- spark.network.crypto.keyLength \| 2.2.0 \| SPARK-19139 \| 8f3f73abc1fe62496722476460c174af0250e3fe#diff-0ac65da2bc6b083fb861fe410c7688c2 \| spark.network.crypto.keyFactoryAlgorithm \| 2.2.0 \| SPARK-19139 \| 8f3f73abc1fe62496722476460c174af0250e3fe#diff-0ac65da2bc6b083fb861fe410c7688c2 \| spark.network.crypto.config.* \| 2.2.0 \| SPARK-19139 \| 8f3f73abc1fe62496722476460c174af0250e3fe#diff-0ac65da2bc6b083fb861fe410c7688c2 \| spark.network.crypto.saslFallback \| 2.2.0 \| SPARK-19139 \| 8f3f73abc1fe62496722476460c174af0250e3fe#diff-0ac65da2bc6b083fb861fe410c7688c2 \| spark.authenticate.enableSaslEncryption \| 2.2.0 \| SPARK-19139 \| 8f3f73abc1fe62496722476460c174af0250e3fe#diff-0ac65da2bc6b083fb861fe410c7688c2 \| spark.network.sasl.serverAlwaysEncrypt \| 1.4.0 \| SPARK-6229 \| 38d4e9e446b425ca6a8fe8d8080f387b08683842#diff-d2ce9b38bdc38ca9d7119f9c2cf79907 \| spark.ui.filters \| 1.0.0 \| SPARK-1189 \| 7edbea41b43e0dc11a2de156be220db8b7952d01#diff-f79a5ead735b3d0b34b6b94486918e1c \| spark.acls.enable \| 1.1.0 \| SPARK-1890 and SPARK-1891 \| e3fe6571decfdc406ec6d505fd92f9f2b85a618c#diff-afd88f677ec5ff8b5e96a5cbbe00cd98 \| spark.ui.view.acls \| 1.0.0 \| SPARK-1189 \| 7edbea41b43e0dc11a2de156be220db8b7952d01#diff-afd88f677ec5ff8b5e96a5cbbe00cd98 \| spark.ui.view.acls.groups \| 2.0.0 \| SPARK-4224 \| ae79032dcf160796851ca29116cca146c4d86ada#diff-afd88f677ec5ff8b5e96a5cbbe00cd98 \| spark.admin.acls \| 1.1.0 \| SPARK-1890 and SPARK-1891 \| e3fe6571decfdc406ec6d505fd92f9f2b85a618c#diff-afd88f677ec5ff8b5e96a5cbbe00cd98 \| spark.admin.acls.groups \| 2.0.0 \| SPARK-4224 \| ae79032dcf160796851ca29116cca146c4d86ada#diff-afd88f677ec5ff8b5e96a5cbbe00cd98 \| spark.modify.acls \| 1.1.0 \| SPARK-1890 and SPARK-1891 \| e3fe6571decfdc406ec6d505fd92f9f2b85a618c#diff-afd88f677ec5ff8b5e96a5cbbe00cd98 \| spark.modify.acls.groups \| 2.0.0 \| SPARK-4224 \| ae79032dcf160796851ca29116cca146c4d86ada#diff-afd88f677ec5ff8b5e96a5cbbe00cd98 \| spark.user.groups.mapping \| 2.0.0 \| SPARK-4224 \| ae79032dcf160796851ca29116cca146c4d86ada#diff-afd88f677ec5ff8b5e96a5cbbe00cd98 \| spark.history.ui.acls.enable \| 1.0.1 \| Spark 1489 \| c8dd13221215275948b1a6913192d40e0c8cbadd#diff-b49b5b9c31ddb36a9061004b5b723058 \| spark.history.ui.admin.acls \| 2.1.1 \| SPARK-19033 \| 4ca1788805e4a0131ba8f0ccb7499ee0e0242837#diff-a7befb99e7bd7e3ab5c46c2568aa5b3e \| spark.history.ui.admin.acls.groups \| 2.1.1 \| SPARK-19033 \| 4ca1788805e4a0131ba8f0ccb7499ee0e0242837#diff-a7befb99e7bd7e3ab5c46c2568aa5b3e \| spark.ui.xXssProtection \| 2.3.0 \| SPARK-22188 \| 5a07aca4d464e96d75ea17bf6768e24b829872ec#diff-6bdad48cfc34314e89599655442ff210 \| spark.ui.xContentTypeOptions.enabled \| 2.3.0 \| SPARK-22188 \| 5a07aca4d464e96d75ea17bf6768e24b829872ec#diff-6bdad48cfc34314e89599655442ff210 \| spark.ui.strictTransportSecurity \| 2.3.0 \| SPARK-22188 \| 5a07aca4d464e96d75ea17bf6768e24b829872ec#diff-6bdad48cfc34314e89599655442ff210 \| spark.security.credentials.${service}.enabled \| 2.3.0 \| SPARK-20434 \| a18d637112b97d2caaca0a8324bdd99086664b24#diff-da6c1fd6d8b0c7538a3e77a09e06a083 \| spark.kerberos.access.hadoopFileSystems \| 3.0.0 \| SPARK-26766 \| d0443a74d185ec72b747fa39994fa9a40ce974cf#diff-6bdad48cfc34314e89599655442ff210 \| ### Why are the changes needed? Supplemental configuration version information. ### Does this PR introduce any user-facing change? 'No'. ### How was this patch tested? Jenkins test Closes #28044 from beliefer/supplement-version-to-security-doc. Authored-by: beliefer <beliefer@163.com> Signed-off-by: HyukjinKwon <gurwls223@apache.org>	2020-03-31 12:33:01 +09:00
beliefer	18b73a5b59	[SPARK-31269][DOC] Supplement version for configuration only appear in configuration doc ### What changes were proposed in this pull request? The `configuration.md` exists some config not organized by `ConfigEntry`. This PR supplements version for configuration only appear in configuration doc. I sorted out some information show below. Item name \| Since version \| JIRA ID \| Commit ID \| Note -- \| -- \| -- \| -- \| -- spark.app.name \| 0.9.0 \| None \| 994f080f8ae3372366e6004600ba791c8a372ff0#diff-529fc5c06b9731c1fbda6f3db60b16aa \| spark.driver.resource.{resourceName}.amount \| 3.0.0 \| SPARK-27760 \| d30284b5a51dd784f663eb4eea37087b35a54d00#diff-76e731333fb756df3bff5ddb3b731c46 \| spark.driver.resource.{resourceName}.discoveryScript \| 3.0.0 \| SPARK-27488 \| 74e5e41eebf9ed596b48e6db52a2a9c642e5cbc3#diff-76e731333fb756df3bff5ddb3b731c46 \| spark.driver.resource.{resourceName}.vendor \| 3.0.0 \| SPARK-27362 \| 1277f8fa92da85d9e39d9146e3099fcb75c71a8f#diff-76e731333fb756df3bff5ddb3b731c46 \| spark.executor.resource.{resourceName}.amount \| 3.0.0 \| SPARK-27760 \| d30284b5a51dd784f663eb4eea37087b35a54d00#diff-76e731333fb756df3bff5ddb3b731c46 \| spark.executor.resource.{resourceType}.discoveryScript \| 3.0.0 \| SPARK-27024 \| db2e3c43412e4a7fb4a46c58d73d9ab304a1e949#diff-76e731333fb756df3bff5ddb3b731c46 \| spark.executor.resource.{resourceName}.vendor \| 3.0.0 \| SPARK-27362 \| 1277f8fa92da85d9e39d9146e3099fcb75c71a8f#diff-76e731333fb756df3bff5ddb3b731c46 \| spark.local.dir \| 0.5.0 \| None \| 0e93891d3d7df849cff6442038c111ffd42a5243#diff-17fd275d280b667722664ed833c6402a \| spark.logConf \| 0.9.0 \| None \| d8bcc8e9a095c1b20dd7a17b6535800d39bff80e#diff-364713d7776956cb8b0a771e9b62f82d \| spark.master \| 0.9.0 \| SPARK-544 \| 2573add94cf920a88f74d80d8ea94218d812704d#diff-529fc5c06b9731c1fbda6f3db60b16aa \| spark.driver.defaultJavaOptions \| 3.0.0 \| SPARK-23472 \| f83000597f250868de9722d8285fed013abc5ecf#diff-a78ecfc6a89edfaf0b60a5eaa0381970 \| spark.executor.defaultJavaOptions \| 3.0.0 \| SPARK-23472 \| f83000597f250868de9722d8285fed013abc5ecf#diff-a78ecfc6a89edfaf0b60a5eaa0381970 \| spark.executorEnv.[EnvironmentVariableName] \| 0.9.0 \| None \| 642029e7f43322f84abe4f7f36bb0b1b95d8101d#diff-529fc5c06b9731c1fbda6f3db60b16aa \| spark.python.profile \| 1.2.0 \| SPARK-3478 \| 1aa549ba9839565274a12c52fa1075b424f138a6#diff-d6fe2792e44f6babc94aabfefc8b9bce \| spark.python.profile.dump \| 1.2.0 \| SPARK-3478 \| 1aa549ba9839565274a12c52fa1075b424f138a6#diff-d6fe2792e44f6babc94aabfefc8b9bce \| spark.python.worker.memory \| 1.1.0 \| SPARK-2538 \| 14174abd421318e71c16edd24224fd5094bdfed4#diff-d6fe2792e44f6babc94aabfefc8b9bce \| spark.jars.packages \| 1.5.0 \| SPARK-9263 \| 34335719a372c1951fdb4dd25b75b086faf1076f#diff-63a5d817d2d45ae24de577f6a1bd80f9 \| spark.jars.excludes \| 1.5.0 \| SPARK-9263 \| 34335719a372c1951fdb4dd25b75b086faf1076f#diff-63a5d817d2d45ae24de577f6a1bd80f9 \| spark.jars.ivy \| 1.3.0 \| SPARK-5341 \| 3b7acd22ab4a134c74746e3b9a803dbd34d43855#diff-63a5d817d2d45ae24de577f6a1bd80f9 \| spark.jars.ivySettings \| 2.2.0 \| SPARK-17568 \| 3bc2eff8880a3ba8d4318118715ea1a47048e3de#diff-4d2ab44195558d5a9d5f15b8803ef39d \| spark.jars.repositories \| 2.3.0 \| SPARK-21403 \| d8257b99ddae23f702f312640a5335ddb4554403#diff-4d2ab44195558d5a9d5f15b8803ef39d \| spark.shuffle.io.maxRetries \| 1.2.0 \| SPARK-4188 \| c1ea5c542f3267c0b23a7775887e3a6ece793fe3#diff-d2ce9b38bdc38ca9d7119f9c2cf79907 \| spark.shuffle.io.numConnectionsPerPeer \| 1.2.1 \| SPARK-4740 \| 441ec3451730c7ae3dbef8952e313071d6147ab6#diff-d2ce9b38bdc38ca9d7119f9c2cf79907 \| spark.shuffle.io.preferDirectBufs \| 1.2.0 \| SPARK-4188 \| c1ea5c542f3267c0b23a7775887e3a6ece793fe3#diff-d2ce9b38bdc38ca9d7119f9c2cf79907 \| spark.shuffle.io.retryWait \| 1.2.1 \| None \| 5e5d8f469a1bea9bbe606f772ccdcab7c184c651#diff-d2ce9b38bdc38ca9d7119f9c2cf79907 \| spark.shuffle.io.backLog \| 1.1.1 \| SPARK-2468 \| 66b4c81db7e826c00f7fb449b8a8af810cf7dd9a#diff-bdee8e601924d41e93baa7287189e878 \| spark.shuffle.service.index.cache.size \| 2.3.0 \| SPARK-21501 \| 1662e93119d68498942386906de309d35f4a135f#diff-97d5edc927a83a678e013ae00343df94 \| spark.shuffle.maxChunksBeingTransferred \| 2.3.0 \| SPARK-21175 \| 799e13161e89f1ea96cb1bc7b507a05af2e89cd0#diff-0ac65da2bc6b083fb861fe410c7688c2 \| spark.sql.ui.retainedExecutions \| 1.5.0 \| SPARK-8861 and SPARK-8862 \| ebc3aad272b91cf58e2e1b4aa92b49b8a947a045#diff-81764e4d52817f83bdd5336ef1226bd9 \| spark.streaming.ui.retainedBatches \| 1.0.0 \| SPARK-1386 \| f36dc3fed0a0671b0712d664db859da28c0a98e2#diff-56b8d67d07284cfab165d5363bd3500e \| spark.default.parallelism \| 0.5.0 \| None \| e5c4cd8a5e188592f8786a265c0cd073c69ac886#diff-0544ebf7533fa70ff5103e0fe1f0b036 \| spark.files.fetchTimeout \| 1.0.0 \| None \| f6f9d02e85d17da2f742ed0062f1648a9293e73c#diff-d239aee594001f8391676e1047a0381e \| spark.files.useFetchCache \| 1.2.2 \| SPARK-6313 \| a2a94a154bdd00753b8d5e344d712664c7151050#diff-d239aee594001f8391676e1047a0381e \| spark.files.overwrite \| 1.0.0 \| None \| 84670f2715392859624df290c1b52eb4ed4a9cb1#diff-d239aee594001f8391676e1047a0381e \| Exists in branch-1.0, but the version of pom is 0.9.0-incubating-SNAPSHOT spark.hadoop.cloneConf \| 1.0.3 \| SPARK-2546 \| 6d8f1dd15afdc7432b5721c89f9b2b402460322b#diff-83eb37f7b0ebed3c14ccb7bff0d577c2 \| spark.hadoop.validateOutputSpecs \| 1.0.1 \| SPARK-1677 \| 8100cbdb7546e8438019443cfc00683017c81278#diff-f70e97c099b5eac05c75288cb215e080 \| spark.hadoop.mapreduce.fileoutputcommitter.algorithm.version \| 2.2.0 \| SPARK-20107 \| edc87d76efea7b4d19d9d0c4ddba274a3ccb8752#diff-76e731333fb756df3bff5ddb3b731c46 \| spark.rpc.io.backLog \| 3.0.0 \| SPARK-27868 \| 09ed64d795d3199a94e175273fff6fcea6b52131#diff-76e731333fb756df3bff5ddb3b731c46 \| spark.network.io.preferDirectBufs \| 3.0.0 \| SPARK-24920 \| e103c4a5e72bab8862ff49d6d4c1e62e642fc412#diff-0ac65da2bc6b083fb861fe410c7688c2 \| spark.port.maxRetries \| 1.1.1 \| SPARK-3565 \| 32f2222e915f31422089139944a077e2cbd442f9#diff-d239aee594001f8391676e1047a0381e \| spark.core.connection.ack.wait.timeout \| 1.1.1 \| SPARK-2677 \| bd3ce2ffb8964abb4d59918ebb2c230fe4614aa2#diff-f748e95f2aa97ed715afa53ddeeac9de \| spark.scheduler.listenerbus.eventqueue.shared.capacity \| 3.0.0 \| SPARK-28574 \| c212c9d9ed7375cd1ea16c118733edd84037ec0d#diff-eb519ad78cc3cf0b95839cc37413b509 \| spark.scheduler.listenerbus.eventqueue.appStatus.capacity \| 3.0.0 \| SPARK-28574 \| c212c9d9ed7375cd1ea16c118733edd84037ec0d#diff-eb519ad78cc3cf0b95839cc37413b509 \| spark.scheduler.listenerbus.eventqueue.executorManagement.capacity \| 3.0.0 \| SPARK-28574 \| c212c9d9ed7375cd1ea16c118733edd84037ec0d#diff-eb519ad78cc3cf0b95839cc37413b509 \| spark.scheduler.listenerbus.eventqueue.eventLog.capacity \| 3.0.0 \| SPARK-28574 \| c212c9d9ed7375cd1ea16c118733edd84037ec0d#diff-eb519ad78cc3cf0b95839cc37413b509 \| spark.scheduler.listenerbus.eventqueue.streams.capacity \| 3.0.0 \| SPARK-28574 \| c212c9d9ed7375cd1ea16c118733edd84037ec0d#diff-eb519ad78cc3cf0b95839cc37413b509 \| spark.task.resource.{resourceName}.amount \| 3.0.0 \| SPARK-27760 \| d30284b5a51dd784f663eb4eea37087b35a54d00#diff-76e731333fb756df3bff5ddb3b731c46 \| spark.stage.maxConsecutiveAttempts \| 2.2.0 \| SPARK-13369 \| 7b5d873aef672aa0aee41e338bab7428101e1ad3#diff-6a9ff7fb74fd490a50462d45db2d5e11 \| spark.{driver\\|executor}.rpc.io.serverThreads \| 1.6.0 \| SPARK-10745 \| 7c5b641808740ba5eed05ba8204cdbaf3fc579f5#diff-d2ce9b38bdc38ca9d7119f9c2cf79907 \| spark.{driver\\|executor}.rpc.io.clientThreads \| 1.6.0 \| SPARK-10745 \| 7c5b641808740ba5eed05ba8204cdbaf3fc579f5#diff-d2ce9b38bdc38ca9d7119f9c2cf79907 \| spark.{driver\\|executor}.rpc.netty.dispatcher.numThreads \| 3.0.0 \| SPARK-29398 \| 2f0a38cb50e3e8b4b72219c7b2b8b15d51f6b931#diff-a68a21481fea5053848ca666dd3201d8 \| spark.r.driver.command \| 1.5.3 \| SPARK-10971 \| 9695f452e86a88bef3bcbd1f3c0b00ad9e9ac6e1#diff-025470e1b7094d7cf4a78ea353fb3981 \| spark.r.shell.command \| 2.1.0 \| SPARK-17178 \| fa6347938fc1c72ddc03a5f3cd2e929b5694f0a6#diff-a78ecfc6a89edfaf0b60a5eaa0381970 \| spark.graphx.pregel.checkpointInterval \| 2.2.0 \| SPARK-5484 \| f971ce5dd0788fe7f5d2ca820b9ea3db72033ddc#diff-e399679417ffa6eeedf26a7630baca16 \| ### Why are the changes needed? Supplemental configuration version information. ### Does this PR introduce any user-facing change? 'No'. ### How was this patch tested? Jenkins test Closes #28035 from beliefer/supplement-configuration-version. Authored-by: beliefer <beliefer@163.com> Signed-off-by: HyukjinKwon <gurwls223@apache.org>	2020-03-31 12:32:04 +09:00
beliefer	bed21770af	[SPARK-31215][SQL][DOC] Add version information to the static configuration of SQL ### What changes were proposed in this pull request? Add version information to the static configuration of `SQL`. I sorted out some information show below. Item name \| Since version \| JIRA ID \| Commit ID \| Note -- \| -- \| -- \| -- \| -- spark.sql.warehouse.dir \| 2.0.0 \| SPARK-14994 \| 054f991c4350af1350af7a4109ee77f4a34822f0#diff-32bb9518401c0948c5ea19377b5069ab \| spark.sql.catalogImplementation \| 2.0.0 \| SPARK-14720 and SPARK-13643 \| 8fc267ab3322e46db81e725a5cb1adb5a71b2b4d#diff-6bdad48cfc34314e89599655442ff210 \| spark.sql.globalTempDatabase \| 2.1.0 \| SPARK-17338 \| 23ddff4b2b2744c3dc84d928e144c541ad5df376#diff-6bdad48cfc34314e89599655442ff210 \| spark.sql.sources.schemaStringLengthThreshold \| 1.3.1 \| SPARK-6024 \| 6200f0709c5c8440decae8bf700d7859f32ac9d5#diff-41ef65b9ef5b518f77e2a03559893f4d \| 1.3 spark.sql.filesourceTableRelationCacheSize \| 2.2.0 \| SPARK-19265 \| 9d9d67c7957f7cbbdbe889bdbc073568b2bfbb16#diff-32bb9518401c0948c5ea19377b5069ab \| spark.sql.codegen.cache.maxEntries \| 2.4.0 \| SPARK-24727 \| b2deef64f604ddd9502a31105ed47cb63470ec85#diff-5081b9388de3add800b6e4a6ddf55c01 \| spark.sql.codegen.comments \| 2.0.0 \| SPARK-15680 \| f0e8738c1ec0e4c5526aeada6f50cf76428f9afd#diff-8bcc5aea39c73d4bf38aef6f6951d42c \| spark.sql.debug \| 2.1.0 \| SPARK-17899 \| db8784feaa605adcbd37af4bc8b7146479b631f8#diff-32bb9518401c0948c5ea19377b5069ab \| spark.sql.hive.thriftServer.singleSession \| 1.6.0 \| SPARK-11089 \| 167ea61a6a604fd9c0b00122a94d1bc4b1de24ff#diff-ff50aea397a607b79df9bec6f2a841db \| spark.sql.extensions \| 2.2.0 \| SPARK-18127 \| f0de600797ff4883927d0c70732675fd8629e239#diff-5081b9388de3add800b6e4a6ddf55c01 \| spark.sql.queryExecutionListeners \| 2.3.0 \| SPARK-19558 \| bd4eb9ce57da7bacff69d9ed958c94f349b7e6fb#diff-5081b9388de3add800b6e4a6ddf55c01 \| spark.sql.streaming.streamingQueryListeners \| 2.4.0 \| SPARK-24479 \| 7703b46d2843db99e28110c4c7ccf60934412504#diff-5081b9388de3add800b6e4a6ddf55c01 \| spark.sql.ui.retainedExecutions \| 1.5.0 \| SPARK-8861 and SPARK-8862 \| ebc3aad272b91cf58e2e1b4aa92b49b8a947a045#diff-81764e4d52817f83bdd5336ef1226bd9 \| spark.sql.broadcastExchange.maxThreadThreshold \| 3.0.0 \| SPARK-26601 \| 126310ca68f2f248ea8b312c4637eccaba2fdc2b#diff-5081b9388de3add800b6e4a6ddf55c01 \| spark.sql.subquery.maxThreadThreshold \| 2.4.6 \| SPARK-30556 \| 2fc562cafd71ec8f438f37a28b65118906ab2ad2#diff-5081b9388de3add800b6e4a6ddf55c01 \| spark.sql.event.truncate.length \| 3.0.0 \| SPARK-27045 \| e60d8fce0b0cf2a6d766ea2fc5f994546550570a#diff-5081b9388de3add800b6e4a6ddf55c01 \| spark.sql.legacy.sessionInitWithConfigDefaults \| 3.0.0 \| SPARK-27253 \| 83f628b57da39ad9732d1393aebac373634a2eb9#diff-5081b9388de3add800b6e4a6ddf55c01 \| spark.sql.defaultUrlStreamHandlerFactory.enabled \| 3.0.0 \| SPARK-25694 \| 8469614c0513fbed87977d4e741649db3fdd8add#diff-5081b9388de3add800b6e4a6ddf55c01 \| spark.sql.streaming.ui.enabled \| 3.0.0 \| SPARK-29543 \| f9b86370cb04b72a4f00cbd4d60873960aa2792c#diff-5081b9388de3add800b6e4a6ddf55c01 \| spark.sql.streaming.ui.retainedProgressUpdates \| 3.0.0 \| SPARK-29543 \| f9b86370cb04b72a4f00cbd4d60873960aa2792c#diff-5081b9388de3add800b6e4a6ddf55c01 \| spark.sql.streaming.ui.retainedQueries \| 3.0.0 \| SPARK-29543 \| f9b86370cb04b72a4f00cbd4d60873960aa2792c#diff-5081b9388de3add800b6e4a6ddf55c01 \| ### Why are the changes needed? Supplemental configuration version information. ### Does this PR introduce any user-facing change? 'No'. ### How was this patch tested? Exists UT Closes #27981 from beliefer/add-version-to-sql-static-config. Authored-by: beliefer <beliefer@163.com> Signed-off-by: HyukjinKwon <gurwls223@apache.org>	2020-03-31 12:31:25 +09:00
zhengruifeng	1dce6c1fd4	[SPARK-31222][ML] Make ANOVATest Sparsity-Aware ### What changes were proposed in this pull request? when input dataset is sparse, make `ANOVATest` only process non-zero value ### Why are the changes needed? for performance ### Does this PR introduce any user-facing change? No ### How was this patch tested? existing testsuites Closes #27982 from zhengruifeng/anova_sparse. Authored-by: zhengruifeng <ruifengz@foxmail.com> Signed-off-by: zhengruifeng <ruifengz@foxmail.com>	2020-03-31 10:40:17 +08:00
Dongjoon Hyun	cda2e30e77	Revert "[SPARK-31280][SQL] Perform propagating empty relation after RewritePredicateSubquery" This reverts commit `f376d24ea1`.	2020-03-30 19:14:14 -07:00
Luca Canali	aa98ac52db	[SPARK-30775][DOC] Improve the description of executor metrics in the monitoring documentation ### What changes were proposed in this pull request? This PR (SPARK-30775) aims to improve the description of the executor metrics in the monitoring documentation. ### Why are the changes needed? Improve and clarify monitoring documentation by: - adding reference to the Prometheus end point, as implemented in [SPARK-29064] - extending the list and descripion of executor metrics, following up from [SPARK-27157] ### Does this PR introduce any user-facing change? Documentation update. ### How was this patch tested? n.a. Closes #27526 from LucaCanali/docPrometheusMetricsFollowupSpark29064. Authored-by: Luca Canali <luca.canali@cern.ch> Signed-off-by: Dongjoon Hyun <dongjoon@apache.org>	2020-03-30 18:00:54 -07:00
Đặng Minh Dũng	1d0fc9aa85	[SPARK-29574][K8S][FOLLOWUP] Fix bash comparison error in Docker entrypoint.sh ### What changes were proposed in this pull request? A small change to fix an error in Docker `entrypoint.sh` ### Why are the changes needed? When spark running on Kubernetes, I got the following logs: ```log + '[' -n ']' + '[' -z ']' ++ /bin/hadoop classpath /opt/entrypoint.sh: line 62: /bin/hadoop: No such file or directory + export SPARK_DIST_CLASSPATH= + SPARK_DIST_CLASSPATH= ``` This is because you are missing some quotes on bash comparisons. ### Does this PR introduce any user-facing change? No ## How was this patch tested? CI Closes #28075 from dungdm93/patch-1. Authored-by: Đặng Minh Dũng <dungdm93@live.com> Signed-off-by: Dongjoon Hyun <dongjoon@apache.org>	2020-03-30 15:41:57 -07:00
manuzhang	0d997e5156	[SPARK-31219][YARN] Enable closeIdleConnections in YarnShuffleService ### What changes were proposed in this pull request? Close idle connections at shuffle server side when an `IdleStateEvent` is triggered after `spark.shuffle.io.connectionTimeout` or `spark.network.timeout` time. It's based on following investigations. 1. We found connections on our clusters building up continuously (> 10k for some nodes). Is that normal ? We don't think so. 2. We looked into the connections on one node and found there were a lot of half-open connections. (connections only existed on one node) 3. We also checked those connections were very old (> 21 hours). (FYI, https://superuser.com/questions/565991/how-to-determine-the-socket-connection-up-time-on-linux) 4. Looking at the code, TransportContext registers an IdleStateHandler which should fire an IdleStateEvent when timeout. We did a heap dump of the YarnShuffleService and checked the attributes of IdleStateHandler. It turned out firstAllIdleEvent of many IdleStateHandlers were already false so IdleStateEvent were already fired. 5. Finally, we realized the IdleStateEvent would not be handled since closeIdleConnections are hardcoded to false for YarnShuffleService. ### Why are the changes needed? Idle connections to YarnShuffleService could never be closed, and will be accumulating and taking up memory and file descriptors. ### Does this PR introduce any user-facing change? No. ### How was this patch tested? Existing tests. Closes #27998 from manuzhang/spark-31219. Authored-by: manuzhang <owenzhang1990@gmail.com> Signed-off-by: Thomas Graves <tgraves@apache.org>	2020-03-30 12:44:46 -05:00
Maxim Gekk	a1dbcd13a3	[SPARK-31296][SQL][TESTS] Benchmark date-time rebasing in Parquet datasource ### What changes were proposed in this pull request? In the PR, I propose to add new benchmark `DateTimeRebaseBenchmark` which should measure the performance of rebasing of dates/timestamps from/to to the hybrid calendar (Julian+Gregorian) to/from Proleptic Gregorian calendar: 1. In write, it saves separately dates and timestamps before and after 1582 year w/ and w/o rebasing. 2. In read, it loads previously saved parquet files by vectorized reader and by regular reader. Here is the summary of benchmarking: - Saving timestamps is ~6 times slower - Loading timestamps w/ vectorized off is ~4 times slower - Loading timestamps w/ vectorized on is ~10 times slower ### Why are the changes needed? To know the impact of date-time rebasing introduced by #27915, #27953, #27807. ### Does this PR introduce any user-facing change? No ### How was this patch tested? Run the `DateTimeRebaseBenchmark` benchmark using Amazon EC2: \| Item \| Description \| \| ---- \| ----\| \| Region \| us-west-2 (Oregon) \| \| Instance \| r3.xlarge \| \| AMI \| ubuntu/images/hvm-ssd/ubuntu-bionic-18.04-amd64-server-20190722.1 (ami-06f2f779464715dc5) \| \| Java \| OpenJDK8/11 \| Closes #28057 from MaxGekk/rebase-bechmark. Lead-authored-by: Maxim Gekk <max.gekk@gmail.com> Co-authored-by: Max Gekk <max.gekk@gmail.com> Signed-off-by: Wenchen Fan <wenchen@databricks.com>	2020-03-30 16:46:31 +08:00
zhengruifeng	34d6b90449	[SPARK-31283][ML] Simplify ChiSq by adding a common method ### What changes were proposed in this pull request? add a common method `computeChiSq` and reuse it in both `chiSquaredDenseFeatures` and `chiSquaredSparseFeatures` ### Why are the changes needed? to simplify ChiSq ### Does this PR introduce any user-facing change? No ### How was this patch tested? existing testsuites Closes #28045 from zhengruifeng/simplify_chisq. Authored-by: zhengruifeng <ruifengz@foxmail.com> Signed-off-by: zhengruifeng <ruifengz@foxmail.com>	2020-03-30 13:37:56 +08:00
Oleksii Kachaiev	22bb6b0fdd	[SPARK-30532] DataFrameStatFunctions to work with TABLE.COLUMN syntax ### What changes were proposed in this pull request? `DataFrameStatFunctions` now works correctly with fully qualified column name (Table.Column syntax) by properly resolving the name instead of relying on field names from schema, notably: * `approxQuantile` * `freqItems` * `cov` * `corr` (other functions from `DataFrameStatFunctions` already work correctly). See code examples below. ### Why are the changes needed? With current implementation some stat functions are impossible to use when joining datasets with similar column names. ### Does this PR introduce any user-facing change? Yes. Before the change, the following code would fail with `AnalysisException`. ```scala scala> val df1 = sc.parallelize(0 to 10).toDF("num").as("table1") df1: org.apache.spark.sql.Dataset[org.apache.spark.sql.Row] = [num: int] scala> val df2 = sc.parallelize(0 to 10).toDF("num").as("table2") df2: org.apache.spark.sql.Dataset[org.apache.spark.sql.Row] = [num: int] scala> val dfx = df2.crossJoin(df1) dfx: org.apache.spark.sql.DataFrame = [num: int, num: int] scala> dfx.stat.approxQuantile("table1.num", Array(0.1), 0.0) res0: Array[Double] = Array(1.0) scala> dfx.stat.corr("table1.num", "table2.num") res1: Double = 1.0 scala> dfx.stat.cov("table1.num", "table2.num") res2: Double = 11.0 scala> dfx.stat.freqItems(Array("table1.num", "table2.num")) res3: org.apache.spark.sql.DataFrame = [table1.num_freqItems: array<int>, table2.num_freqItems: array<int>] ``` ### How was this patch tested? Corresponding unit tests are added to `DataFrameStatSuite.scala` (marked as "SPARK-30532"). Closes #27916 from kachayev/fix-spark-30532. Authored-by: Oleksii Kachaiev <kachayev@gmail.com> Signed-off-by: Wenchen Fan <wenchen@databricks.com>	2020-03-30 13:20:57 +08:00
Maxim Gekk	d2ff5c5bfb	[SPARK-31286][SQL][DOC] Specify formats of time zone ID for JSON/CSV option and from/to_utc_timestamp ### What changes were proposed in this pull request? In the PR, I propose to update the doc for the `timeZone` option in JSON/CSV datasources and for the `tz` parameter of the `from_utc_timestamp()`/`to_utc_timestamp()` functions, and to restrict format of config's values to 2 forms: 1. Geographical regions, such as `America/Los_Angeles`. 2. Fixed offsets - a fully resolved offset from UTC. For example, `-08:00`. ### Why are the changes needed? Other formats such as three-letter time zone IDs are ambitious, and depend on the locale. For example, `CST` could be U.S. `Central Standard Time` and `China Standard Time`. Such formats have been already deprecated in JDK, see [Three-letter time zone IDs](https://docs.oracle.com/javase/8/docs/api/java/util/TimeZone.html). ### Does this PR introduce any user-facing change? No ### How was this patch tested? By running `./dev/scalastyle`, and manual testing. Closes #28051 from MaxGekk/doc-time-zone-option. Authored-by: Maxim Gekk <max.gekk@gmail.com> Signed-off-by: Wenchen Fan <wenchen@databricks.com>	2020-03-30 12:20:11 +08:00
Kengo Seki	60dd1a690f	[SPARK-31293][DSTREAMS][KINESIS][DOC] Fix wrong examples and help messages for Kinesis integration ### What changes were proposed in this pull request? This PR (SPARK-31293) fixes wrong command examples, parameter descriptions and help message format for Amazon Kinesis integration with Spark Streaming. ### Why are the changes needed? To improve usability of those commands. ### Does this PR introduce any user-facing change? No ### How was this patch tested? I ran the fixed commands manually and confirmed they worked as expected. Closes #28063 from sekikn/SPARK-31293. Authored-by: Kengo Seki <sekikn@apache.org> Signed-off-by: Dongjoon Hyun <dongjoon@apache.org>	2020-03-29 14:27:19 -07:00
Kent Yao	f376d24ea1	[SPARK-31280][SQL] Perform propagating empty relation after RewritePredicateSubquery ### What changes were proposed in this pull request? ```sql scala> spark.sql(" select * from values(1), (2) t(key) where key in (select 1 as key where 1=0)").queryExecution res15: org.apache.spark.sql.execution.QueryExecution = == Parsed Logical Plan == 'Project [] +- 'Filter 'key IN (list#39 []) : +- Project [1 AS key#38] : +- Filter (1 = 0) : +- OneRowRelation +- 'SubqueryAlias t +- 'UnresolvedInlineTable [key], [List(1), List(2)] == Analyzed Logical Plan == key: int Project [key#40] +- Filter key#40 IN (list#39 []) : +- Project [1 AS key#38] : +- Filter (1 = 0) : +- OneRowRelation +- SubqueryAlias t +- LocalRelation [key#40] == Optimized Logical Plan == Join LeftSemi, (key#40 = key#38) :- LocalRelation [key#40] +- LocalRelation <empty>, [key#38] == Physical Plan == (1) BroadcastHashJoin [key#40], [key#38], LeftSemi, BuildRight :- *(1) LocalTableScan [key#40] +- Br... ``` `LocalRelation <empty> ` should be able to propagate after subqueries are lift up to joins ### Why are the changes needed? optimize query ### Does this PR introduce any user-facing change? no ### How was this patch tested? add new tests Closes #28043 from yaooqinn/SPARK-31280. Authored-by: Kent Yao <yaooqinn@hotmail.com> Signed-off-by: Dongjoon Hyun <dongjoon@apache.org>	2020-03-29 11:32:22 -07:00
Huaxin Gao	e656e99061	[SPARK-30363][SQL][DOCS][FOLLOWUP] Fix a broken link in SQL Reference ### What changes were proposed in this pull request? Fix a broken link and make the relevant docs reference to the new doc ### Why are the changes needed? ### Does this PR introduce any user-facing change? Yes, make CACHE TABLE, UNCACHE TABLE, CLEAR CACHE, REFRESH TABLE link to the new doc ### How was this patch tested? Manually build and check Closes #28065 from huaxingao/spark-30363-follow-up. Authored-by: Huaxin Gao <huaxing@us.ibm.com> Signed-off-by: Sean Owen <srowen@gmail.com>	2020-03-29 11:19:24 -05:00
gatorsmile	3884455780	[SPARK-31087] [SQL] Add Back Multiple Removed APIs ### What changes were proposed in this pull request? Based on the discussion in the mailing list [[Proposal] Modification to Spark's Semantic Versioning Policy](http://apache-spark-developers-list.1001551.n3.nabble.com/Proposal-Modification-to-Spark-s-Semantic-Versioning-Policy-td28938.html) , this PR is to add back the following APIs whose maintenance cost are relatively small. - functions.toDegrees/toRadians - functions.approxCountDistinct - functions.monotonicallyIncreasingId - Column.!== - Dataset.explode - Dataset.registerTempTable - SQLContext.getOrCreate, setActive, clearActive, constructors Below is the other removed APIs in the original PR, but not added back in this PR [https://issues.apache.org/jira/browse/SPARK-25908]: - Remove some AccumulableInfo .apply() methods - Remove non-label-specific multiclass precision/recall/fScore in favor of accuracy - Remove unused Python StorageLevel constants - Remove unused multiclass option in libsvm parsing - Remove references to deprecated spark configs like spark.yarn.am.port - Remove TaskContext.isRunningLocally - Remove ShuffleMetrics.shuffle* methods - Remove BaseReadWrite.context in favor of session ### Why are the changes needed? Avoid breaking the APIs that are commonly used. ### Does this PR introduce any user-facing change? Adding back the APIs that were removed in 3.0 branch does not introduce the user-facing changes, because Spark 3.0 has not been released. ### How was this patch tested? Added a new test suite for these APIs. Author: gatorsmile <gatorsmile@gmail.com> Author: yi.wu <yi.wu@databricks.com> Closes #27821 from gatorsmile/addAPIBackV2.	2020-03-28 22:05:16 -07:00
HyukjinKwon	3165a95a04	[SPARK-31287][PYTHON][SQL] Ignore type hints in groupby.(cogroup.)applyInPandas and mapInPandas ### What changes were proposed in this pull request? This PR proposes to make pandas function APIs (`groupby.(cogroup.)applyInPandas` and `mapInPandas`) to ignore Python type hints. ### Why are the changes needed? Python type hints are optional. It shouldn't affect where pandas UDFs are not used. This is also a future work for them to support other type hints. We shouldn't at least throw an exception at this moment. ### Does this PR introduce any user-facing change? No, it's master-only change. ```python import pandas as pd def pandas_plus_one(pdf: pd.DataFrame) -> pd.DataFrame: return pdf + 1 spark.range(10).groupby('id').applyInPandas(pandas_plus_one, schema="id long").show() ``` ```python import pandas as pd def pandas_plus_one(left: pd.DataFrame, right: pd.DataFrame) -> pd.DataFrame: return left + 1 spark.range(10).groupby('id').cogroup(spark.range(10).groupby("id")).applyInPandas(pandas_plus_one, schema="id long").show() ``` ```python from typing import Iterator import pandas as pd def pandas_plus_one(iter: Iterator[pd.DataFrame]) -> Iterator[pd.DataFrame]: return map(lambda v: v + 1, iter) spark.range(10).mapInPandas(pandas_plus_one, schema="id long").show() ``` Before: Exception After: ``` +---+ \| id\| +---+ \| 1\| \| 2\| \| 3\| \| 4\| \| 5\| \| 6\| \| 7\| \| 8\| \| 9\| \| 10\| +---+ ``` ### How was this patch tested? Closes #28052 from HyukjinKwon/SPARK-31287. Authored-by: HyukjinKwon <gurwls223@apache.org> Signed-off-by: HyukjinKwon <gurwls223@apache.org>	2020-03-29 13:59:18 +09:00
Zhenhua Wang	791d2ba346	[SPARK-31261][SQL] Avoid npe when reading bad csv input with `columnNameCorruptRecord` specified ### What changes were proposed in this pull request? SPARK-25387 avoids npe for bad csv input, but when reading bad csv input with `columnNameCorruptRecord` specified, `getCurrentInput` is called and it still throws npe. ### Why are the changes needed? Bug fix. ### Does this PR introduce any user-facing change? No. ### How was this patch tested? Add a test. Closes #28029 from wzhfy/corrupt_column_npe. Authored-by: Zhenhua Wang <wzh_zju@163.com> Signed-off-by: HyukjinKwon <gurwls223@apache.org>	2020-03-29 13:30:14 +09:00
HyukjinKwon	34c7476cb5	[SPARK-30722][DOCS][FOLLOW-UP] Add Pandas Function API into the menu ### What changes were proposed in this pull request? This PR adds "Pandas Function API" into the menu. ### Why are the changes needed? To be consistent and to make easier to navigate. ### Does this PR introduce any user-facing change? No, master only. ![Screen Shot 2020-03-27 at 11 40 29 PM](https://user-images.githubusercontent.com/6477701/77767405-60306600-7084-11ea-944a-93726259cd00.png) ### How was this patch tested? Manually verified by `SKIP_API=1 jekyll build`. Closes #28054 from HyukjinKwon/followup-spark-30722. Authored-by: HyukjinKwon <gurwls223@apache.org> Signed-off-by: Dongjoon Hyun <dongjoon@apache.org>	2020-03-28 18:36:34 -07:00
Kengo Seki	0b237bd615	[SPARK-31292][CORE][SQL] Replace toSet.toSeq with distinct for readability ### What changes were proposed in this pull request? This PR replaces the method calls of `toSet.toSeq` with `distinct`. ### Why are the changes needed? `toSet.toSeq` is intended to make its elements unique but a bit verbose. Using `distinct` instead is easier to understand and improves readability. ### Does this PR introduce any user-facing change? No ### How was this patch tested? Tested with the existing unit tests and found no problem. Closes #28062 from sekikn/SPARK-31292. Authored-by: Kengo Seki <sekikn@apache.org> Signed-off-by: Takeshi Yamamuro <yamamuro@apache.org>	2020-03-29 08:48:08 +09:00
Dongjoon Hyun	d025ddbaa7	[SPARK-31238][SPARK-31284][TEST][FOLLOWUP] Fix readResourceOrcFile to create a local file from resource ### What changes were proposed in this pull request? This PR aims to copy a test resource file to a local file in `OrcTest` suite before reading it. ### Why are the changes needed? SPARK-31238 and SPARK-31284 added test cases to access the resouce file in `sql/core` module from `sql/hive` module. In Maven test environment, this causes a failure. ``` - SPARK-31238: compatibility with Spark 2.4 in reading dates * FAILED * java.lang.IllegalArgumentException: java.net.URISyntaxException: Relative path in absolute URI: jar:file:/home/jenkins/workspace/spark-master-test-maven-hadoop-3.2-hive-2.3-jdk-11/sql/core/target/spark-sql_2.12-3.1.0-SNAPSHOT-tests.jar!/test-data/before_1582_date_v2_4.snappy.orc ``` ``` - SPARK-31284: compatibility with Spark 2.4 in reading timestamps * FAILED * java.lang.IllegalArgumentException: java.net.URISyntaxException: Relative path in absolute URI: jar:file:/home/jenkins/workspace/spark-master-test-maven-hadoop-3.2-hive-2.3/sql/core/target/spark-sql_2.12-3.1.0-SNAPSHOT-tests.jar!/test-data/before_1582_ts_v2_4.snappy.orc ``` ### Does this PR introduce any user-facing change? No ### How was this patch tested? Pass the Jenkins with Maven. Closes #28059 from dongjoon-hyun/SPARK-31238. Authored-by: Dongjoon Hyun <dongjoon@apache.org> Signed-off-by: Dongjoon Hyun <dongjoon@apache.org>	2020-03-27 18:44:53 -07:00
Wenchen Fan	c4e98c065c	[SPARK-31271][UI] fix web ui for driver side SQL metrics ### What changes were proposed in this pull request? In https://github.com/apache/spark/pull/23551, we changed the metrics type of driver-side SQL metrics to size/time etc. which comes with max/min/median info. This doesn't make sense for driver side SQL metrics as they have only one value. It makes the web UI hard to read: ![image](https://user-images.githubusercontent.com/3182036/77653892-42db9900-6fab-11ea-8e7f-92f763fa32ff.png) This PR updates the SQL metrics UI to only display max/min/median if there are more than one metrics values: ![image](https://user-images.githubusercontent.com/3182036/77653975-5f77d100-6fab-11ea-849e-64c935377c8e.png) ### Why are the changes needed? Makes the UI easier to read ### Does this PR introduce any user-facing change? no ### How was this patch tested? manual test Closes #28037 from cloud-fan/ui. Authored-by: Wenchen Fan <wenchen@databricks.com> Signed-off-by: Dongjoon Hyun <dongjoon@apache.org>	2020-03-27 15:45:35 -07:00
Liang-Chi Hsieh	aa8776bb59	[SPARK-29721][SQL] Prune unnecessary nested fields from Generate without Project ### What changes were proposed in this pull request? This patch proposes to prune unnecessary nested fields from Generate which has no Project on top of it. ### Why are the changes needed? In Optimizer, we can prune nested columns from Project(projectList, Generate). However, unnecessary columns could still possibly be read in Generate, if no Project on top of it. We should prune it too. ### Does this PR introduce any user-facing change? No ### How was this patch tested? Unit test. Closes #27517 from viirya/SPARK-29721-2. Lead-authored-by: Liang-Chi Hsieh <liangchi@uber.com> Co-authored-by: Liang-Chi Hsieh <viirya@gmail.com> Signed-off-by: Dongjoon Hyun <dongjoon@apache.org>	2020-03-27 10:47:21 -07:00
Wenchen Fan	8a5d49610d	[MINOR][DOC] Refine comments of QueryPlan regarding subquery ### What changes were proposed in this pull request? The query plan of Spark SQL is a mutually recursive structure: QueryPlan -> Expression (PlanExpression) -> QueryPlan, but the transformations do not take this into account. This PR refines the comments of `QueryPlan` to highlight this fact. ### Why are the changes needed? better document. ### Does this PR introduce any user-facing change? no ### How was this patch tested? N/A Closes #28050 from cloud-fan/comment. Authored-by: Wenchen Fan <wenchen@databricks.com> Signed-off-by: Dongjoon Hyun <dongjoon@apache.org>	2020-03-27 09:35:35 -07:00
Prashant Sharma	f87957371d	[SPARK-31200][K8S] Enforce to use `https` in /etc/apt/sources.list …n progress errors. ### What changes were proposed in this pull request? Switching to `https` instead of `http` in the debian mirror urls. ### Why are the changes needed? My ISP was trying to intercept (or trying to serve from cache) the `http` traffic and this was causing a very confusing errors while building the spark image. I thought by posting this, I can help someone save his time and energy, if he encounters the same issue. ``` bash-3.2$ bin/docker-image-tool.sh -r scrapcodes -t v3.1.0-f1cc86 build Sending build context to Docker daemon 203.4MB Step 1/18 : ARG java_image_tag=8-jre-slim Step 2/18 : FROM openjdk:${java_image_tag} ---> 381b20190cf7 Step 3/18 : ARG spark_uid=185 ---> Using cache ---> 65c06f86753c Step 4/18 : RUN set -ex && apt-get update && ln -s /lib /lib64 && apt install -y bash tini libc6 libpam-modules krb5-user libnss3 procps && mkdir -p /opt/spark && mkdir -p /opt/spark/examples && mkdir -p /opt/spark/work-dir && touch /opt/spark/RELEASE && rm /bin/sh && ln -sv /bin/bash /bin/sh && echo "auth required pam_wheel.so use_uid" >> /etc/pam.d/su && chgrp root /etc/passwd && chmod ug+rw /etc/passwd && rm -rf /var/cache/apt/* ---> Running in 96bcbe927d35 + apt-get update Get:1 http://deb.debian.org/debian buster InRelease [122 kB] Get:2 http://deb.debian.org/debian buster-updates InRelease [49.3 kB] Get:3 http://deb.debian.org/debian buster/main amd64 Packages [7907 kB] Err:3 http://deb.debian.org/debian buster/main amd64 Packages File has unexpected size (13217 != 7906744). Mirror sync in progress? [IP: 151.101.10.133 80] Hashes of expected file: - Filesize:7906744 [weak] - SHA256:80ed5d1cc1f31a568b77e4fadfd9e01fa4d65e951243fd2ce29eee14d4b532cc - MD5Sum:80b6d9c1b6630b2234161e42f4040ab3 [weak] Release file created at: Sat, 08 Feb 2020 10:57:10 +0000 Get:5 http://deb.debian.org/debian buster-updates/main amd64 Packages [7380 B] Err:5 http://deb.debian.org/debian buster-updates/main amd64 Packages File has unexpected size (13233 != 7380). Mirror sync in progress? [IP: 151.101.10.133 80] Hashes of expected file: - Filesize:7380 [weak] - SHA256:6af9ea081b6a3da33cfaf76a81978517f65d38e45230089a5612e56f2b6b789d Release file created at: Fri, 20 Mar 2020 02:28:11 +0000 Get:4 http://security-cdn.debian.org/debian-security buster/updates InRelease [65.4 kB] Get:6 http://security-cdn.debian.org/debian-security buster/updates/main amd64 Packages [183 kB] Fetched 419 kB in 1s (327 kB/s) Reading package lists... E: Failed to fetch `80ed5d1cc1` File has unexpected size (13217 != 7906744). Mirror sync in progress? [IP: 151.101.10.133 80] Hashes of expected file: - Filesize:7906744 [weak] - SHA256:80ed5d1cc1f31a568b77e4fadfd9e01fa4d65e951243fd2ce29eee14d4b532cc - MD5Sum:80b6d9c1b6630b2234161e42f4040ab3 [weak] Release file created at: Sat, 08 Feb 2020 10:57:10 +0000 E: Failed to fetch `6af9ea081b` File has unexpected size (13233 != 7380). Mirror sync in progress? [IP: 151.101.10.133 80] Hashes of expected file: - Filesize:7380 [weak] - SHA256:6af9ea081b6a3da33cfaf76a81978517f65d38e45230089a5612e56f2b6b789d Release file created at: Fri, 20 Mar 2020 02:28:11 +0000 E: Some index files failed to download. They have been ignored, or old ones used instead. The command '/bin/sh -c set -ex && apt-get update && ln -s /lib /lib64 && apt install -y bash tini libc6 libpam-modules krb5-user libnss3 procps && mkdir -p /opt/spark && mkdir -p /opt/spark/examples && mkdir -p /opt/spark/work-dir && touch /opt/spark/RELEASE && rm /bin/sh && ln -sv /bin/bash /bin/sh && echo "auth required pam_wheel.so use_uid" >> /etc/pam.d/su && chgrp root /etc/passwd && chmod ug+rw /etc/passwd && rm -rf /var/cache/apt/*' returned a non-zero code: 100 Failed to build Spark JVM Docker image, please refer to Docker build output for details. ``` ### Does this PR introduce any user-facing change? No ### How was this patch tested? Manually by switching to `https` mirrors on the offending ISP (I am already on). Closes #27966 from ScrapCodes/docker-mirror. Authored-by: Prashant Sharma <prashsh1@in.ibm.com> Signed-off-by: Dongjoon Hyun <dongjoon@apache.org>	2020-03-27 09:13:55 -07:00
Maxim Gekk	fc2a974e03	[SPARK-31284][SQL][TESTS] Check rebasing of timestamps in ORC datasource ### What changes were proposed in this pull request? In the PR, I propose 2 tests to check that rebasing of timestamps from/to the hybrid calendar (Julian + Gregorian) to/from Proleptic Gregorian calendar works correctly. 1. The test `compatibility with Spark 2.4 in reading timestamps` load ORC file saved by Spark 2.4.5 via: ```shell $ export TZ="America/Los_Angeles" ``` ```scala scala> spark.conf.set("spark.sql.session.timeZone", "America/Los_Angeles") scala> val df = Seq("1001-01-01 01:02:03.123456").toDF("tsS").select($"tsS".cast("timestamp").as("ts")) df: org.apache.spark.sql.DataFrame = [ts: timestamp] scala> df.write.orc("/Users/maxim/tmp/before_1582/2_4_5_ts_orc") scala> spark.read.orc("/Users/maxim/tmp/before_1582/2_4_5_ts_orc").show(false) +--------------------------+ \|ts \| +--------------------------+ \|1001-01-01 01:02:03.123456\| +--------------------------+ ``` 2. The test `rebasing timestamps in write` is round trip test. Since the previous test confirms correct rebasing of timestamps in read. This test should pass only if rebasing works correctly in write. ### Why are the changes needed? To guarantee that rebasing works correctly for timestamps in ORC datasource. ### Does this PR introduce any user-facing change? No ### How was this patch tested? By running `OrcSourceSuite` for Hive 1.2 and 2.3 via the commands: ``` $ build/sbt -Phive-2.3 "test:testOnly OrcSourceSuite" ``` and ``` $ build/sbt -Phive-1.2 "test:testOnly OrcSourceSuite" ``` Closes #28047 from MaxGekk/rebase-ts-orc-test. Authored-by: Maxim Gekk <max.gekk@gmail.com> Signed-off-by: Dongjoon Hyun <dongjoon@apache.org>	2020-03-27 09:06:59 -07:00
Maxim Gekk	9f0c010a5c	[SPARK-31277][SQL][TESTS] Migrate `DateTimeTestUtils` from `TimeZone` to `ZoneId` ### What changes were proposed in this pull request? In the PR, I propose to change types of `DateTimeTestUtils` values and functions by replacing `java.util.TimeZone` to `java.time.ZoneId`. In particular: 1. Type of `ALL_TIMEZONES` is changed to `Seq[ZoneId]`. 2. Remove `val outstandingTimezones: Seq[TimeZone]`. 3. Change the type of the time zone parameter in `withDefaultTimeZone` to `ZoneId`. 4. Modify affected test suites. ### Why are the changes needed? Currently, Spark SQL's date-time expressions and functions have been already ported on Java 8 time API but tests still use old time APIs. In particular, `DateTimeTestUtils` exposes functions that accept only TimeZone instances. This is inconvenient, and CPU consuming because need to convert TimeZone instances to ZoneId instances via strings (zone ids). ### Does this PR introduce any user-facing change? No ### How was this patch tested? By affected test suites executed by jenkins builds. Closes #28033 from MaxGekk/with-default-time-zone. Authored-by: Maxim Gekk <max.gekk@gmail.com> Signed-off-by: Wenchen Fan <wenchen@databricks.com>	2020-03-27 21:14:25 +08:00
Kent Yao	5945d46c11	[SPARK-31225][SQL] Override sql method of OuterReference ### What changes were proposed in this pull request? OuterReference is one LeafExpression, so it's children is Nil, which makes its SQL representation always be outer(). This makes our explain-command and error msg unclear when OuterReference exists. e.g. ```scala org.apache.spark.sql.AnalysisException: Aggregate/Window/Generate expressions are not valid in where clause of the query. Expression in where clause: [(in.`value` = max(outer()))] Invalid expressions: [max(outer())];; ``` This PR override its `sql` method with its `prettyName` and single argment `e`'s `sql` methond ### Why are the changes needed? improve err message ### Does this PR introduce any user-facing change? yes, the err msg caused by OuterReference has changed ### How was this patch tested? modified ut results Closes #27985 from yaooqinn/SPARK-31225. Authored-by: Kent Yao <yaooqinn@hotmail.com> Signed-off-by: Wenchen Fan <wenchen@databricks.com>	2020-03-27 15:21:19 +08:00
gatorsmile	b9eafcb526	[SPARK-31088][SQL] Add back HiveContext and createExternalTable ### What changes were proposed in this pull request? Based on the discussion in the mailing list [[Proposal] Modification to Spark's Semantic Versioning Policy](http://apache-spark-developers-list.1001551.n3.nabble.com/Proposal-Modification-to-Spark-s-Semantic-Versioning-Policy-td28938.html) , this PR is to add back the following APIs whose maintenance cost are relatively small. - HiveContext - createExternalTable APIs ### Why are the changes needed? Avoid breaking the APIs that are commonly used. ### Does this PR introduce any user-facing change? Adding back the APIs that were removed in 3.0 branch does not introduce the user-facing changes, because Spark 3.0 has not been released. ### How was this patch tested? add a new test suite for createExternalTable APIs. Closes #27815 from gatorsmile/addAPIsBack. Lead-authored-by: gatorsmile <gatorsmile@gmail.com> Co-authored-by: yi.wu <yi.wu@databricks.com> Signed-off-by: gatorsmile <gatorsmile@gmail.com>	2020-03-26 23:51:15 -07:00
gatorsmile	b7e4cc775b	[SPARK-31086][SQL] Add Back the Deprecated SQLContext methods ### What changes were proposed in this pull request? Based on the discussion in the mailing list [[Proposal] Modification to Spark's Semantic Versioning Policy](http://apache-spark-developers-list.1001551.n3.nabble.com/Proposal-Modification-to-Spark-s-Semantic-Versioning-Policy-td28938.html) , this PR is to add back the following APIs whose maintenance cost are relatively small. - SQLContext.applySchema - SQLContext.parquetFile - SQLContext.jsonFile - SQLContext.jsonRDD - SQLContext.load - SQLContext.jdbc ### Why are the changes needed? Avoid breaking the APIs that are commonly used. ### Does this PR introduce any user-facing change? Adding back the APIs that were removed in 3.0 branch does not introduce the user-facing changes, because Spark 3.0 has not been released. ### How was this patch tested? The existing tests. Closes #27839 from gatorsmile/addAPIBackV3. Lead-authored-by: gatorsmile <gatorsmile@gmail.com> Co-authored-by: yi.wu <yi.wu@databricks.com> Signed-off-by: gatorsmile <gatorsmile@gmail.com>	2020-03-26 23:49:24 -07:00
DB Tsai	cb0db21373	[SPARK-25556][SPARK-17636][SPARK-31026][SPARK-31060][SQL][TEST-HIVE1.2] Nested Column Predicate Pushdown for Parquet ### What changes were proposed in this pull request? 1. `DataSourceStrategy.scala` is extended to create `org.apache.spark.sql.sources.Filter` from nested expressions. 2. Translation from nested `org.apache.spark.sql.sources.Filter` to `org.apache.parquet.filter2.predicate.FilterPredicate` is implemented to support nested predicate pushdown for Parquet. ### Why are the changes needed? Better performance for handling nested predicate pushdown. ### Does this PR introduce any user-facing change? No ### How was this patch tested? New tests are added. Closes #27728 from dbtsai/SPARK-17636. Authored-by: DB Tsai <d_tsai@apple.com> Signed-off-by: Wenchen Fan <wenchen@databricks.com>	2020-03-27 14:28:57 +08:00
Huaxin Gao	d279dbf09c	[SPARK-31243][ML][PYSPARK] Add ANOVATest and FValueTest to PySpark ### What changes were proposed in this pull request? Add ANOVATest and FValueTest to PySpark ### Why are the changes needed? Parity between Scala and Python. ### Does this PR introduce any user-facing change? Yes. Python ANOVATest and FValueTest ### How was this patch tested? doctest Closes #28012 from huaxingao/stats-python. Authored-by: Huaxin Gao <huaxing@us.ibm.com> Signed-off-by: zhengruifeng <ruifengz@foxmail.com>	2020-03-27 14:05:49 +08:00
Kousuke Saruta	bc37fdc771	[SPARK-31275][WEBUI] Improve the metrics format in ExecutionPage for StageId ### What changes were proposed in this pull request? In ExecutionPage, metrics format for stageId, attemptId and taskId are displayed like `(stageId (attemptId): taskId)` for now. I changed this format like `(stageId.attemptId taskId)`. ### Why are the changes needed? As cloud-fan suggested [here](https://github.com/apache/spark/pull/27927#discussion_r398591519), `stageId.attemptId` is more standard in Spark. ### Does this PR introduce any user-facing change? Yes. Before applying this change, we can see the UI like as follows. ![with-checked](https://user-images.githubusercontent.com/4736016/77682421-42a6c200-6fda-11ea-92e4-e9f4554adb71.png) And after this change applied, we can like as follows. ![fix-merics-format-with-checked](https://user-images.githubusercontent.com/4736016/77682493-61a55400-6fda-11ea-801f-91a67da698fd.png) ### How was this patch tested? Modified `SQLMetricsSuite` and manual test. Closes #28039 from sarutak/improve-metrics-format. Authored-by: Kousuke Saruta <sarutak@oss.nttdata.com> Signed-off-by: Wenchen Fan <wenchen@databricks.com>	2020-03-27 13:35:28 +08:00
Terry Kim	a97d3b9f4f	[SPARK-31204][SQL] HiveResult compatibility for DatasourceV2 command ### What changes were proposed in this pull request? `HiveResult` performs some conversions for commands to be compatible with Hive output, e.g.: ``` // If it is a describe command for a Hive table, we want to have the output format be similar with Hive. case ExecutedCommandExec(_: DescribeCommandBase) => ... // SHOW TABLES in Hive only output table names, while ours output database, table name, isTemp. case command ExecutedCommandExec(s: ShowTablesCommand) if !s.isExtended => ``` This conversion is needed for DatasourceV2 commands as well and this PR proposes to add the conversion for v2 commands `SHOW TABLES` and `DESCRIBE TABLE`. ### Why are the changes needed? This is a bug where conversion is not applied to v2 commands. ### Does this PR introduce any user-facing change? Yes, now the outputs for v2 commands `SHOW TABLES` and `DESCRIBE TABLE` are compatible with HIVE output. For example, with a table created as: ``` CREATE TABLE testcat.ns.tbl (id bigint COMMENT 'col1') USING foo ``` The output of `SHOW TABLES` has changed from ``` ns table ``` to ``` table ``` And the output of `DESCRIBE TABLE` has changed from ``` id bigint col1 # Partitioning Not partitioned ``` to ``` id bigint col1 # Partitioning Not partitioned ``` ### How was this patch tested? Added unit tests. Closes #28004 from imback82/hive_result. Authored-by: Terry Kim <yuminkim@gmail.com> Signed-off-by: Wenchen Fan <wenchen@databricks.com>	2020-03-27 12:48:14 +08:00
Kent Yao	8be16907c2	[SPARK-31170][SQL] Spark SQL Cli should respect hive-site.xml and spark.sql.warehouse.dir ### What changes were proposed in this pull request? In Spark CLI, we create a hive `CliSessionState` and it does not load the `hive-site.xml`. So the configurations in `hive-site.xml` will not take effects like other spark-hive integration apps. Also, the warehouse directory is not correctly picked. If the `default` database does not exist, the `CliSessionState` will create one during the first time it talks to the metastore. The `Location` of the default DB will be neither the value of `spark.sql.warehousr.dir` nor the user-specified value of `hive.metastore.warehourse.dir`, but the default value of `hive.metastore.warehourse.dir `which will always be `/user/hive/warehouse`. This PR fixes CLiSuite failure with the hive-1.2 profile in https://github.com/apache/spark/pull/27933. In https://github.com/apache/spark/pull/27933, we fix the issue in JIRA by deciding the warehouse dir using all properties from spark conf and Hadoop conf, but properties from `--hiveconf` is not included, they will be applied to the `CliSessionState` instance after it initialized. When this command-line option key is `hive.metastore.warehouse.dir`, the actual warehouse dir is overridden. Because of the logic in Hive for creating the non-existing default database changed, that test passed with `Hive 2.3.6` but failed with `1.2`. So in this PR, Hadoop/Hive configurations are ordered by: ` spark.hive.xxx > spark.hadoop.xxx > --hiveconf xxx > hive-site.xml` througth `ShareState.loadHiveConfFile` before sessionState start ### Why are the changes needed? Bugfix for Spark SQL CLI to pick right confs ### Does this PR introduce any user-facing change? yes, 1. the non-exists default database will be created in the location specified by the users via `spark.sql.warehouse.dir` or `hive.metastore.warehouse.dir`, or the default value of `spark.sql.warehouse.dir` if none of them specified. 2. configurations from `hive-site.xml` will not override command-line options or the properties defined with `spark.hadoo(hive).` prefix in spark conf. ### How was this patch tested? add cli ut Closes #27969 from yaooqinn/SPARK-31170-2. Authored-by: Kent Yao <yaooqinn@hotmail.com> Signed-off-by: Wenchen Fan <wenchen@databricks.com>	2020-03-27 12:05:45 +08:00
Liang-Chi Hsieh	559d3e4051	[SPARK-31186][PYSPARK][SQL] toPandas should not fail on duplicate column names ### What changes were proposed in this pull request? When `toPandas` API works on duplicate column names produced from operators like join, we see the error like: ``` ValueError: The truth value of a Series is ambiguous. Use a.empty, a.bool(), a.item(), a.any() or a.all(). ``` This patch fixes the error in `toPandas` API. ### Why are the changes needed? To make `toPandas` work on dataframe with duplicate column names. ### Does this PR introduce any user-facing change? Yes. Previously calling `toPandas` API on a dataframe with duplicate column names will fail. After this patch, it will produce correct result. ### How was this patch tested? Unit test. Closes #28025 from viirya/SPARK-31186. Authored-by: Liang-Chi Hsieh <viirya@gmail.com> Signed-off-by: HyukjinKwon <gurwls223@apache.org>	2020-03-27 12:10:30 +09:00
beliefer	9e0fee933e	[SPARK-31262][SQL][TESTS] Fix bug tests imported bracketed comments ### What changes were proposed in this pull request? This PR related to https://github.com/apache/spark/pull/27481. If test case A uses `--IMPORT` to import test case B contains bracketed comments, the output can't display bracketed comments in golden files well. The content of `nested-comments.sql` show below: ``` -- This test case just used to test imported bracketed comments. -- the first case of bracketed comment --QUERY-DELIMITER-START /* This is the first example of bracketed comment. SELECT 'ommented out content' AS first; / SELECT 'selected content' AS first; --QUERY-DELIMITER-END ``` The test case `comments.sql` imports `nested-comments.sql` below: `--IMPORT nested-comments.sql` Before this PR, the output will be: ``` -- !query / This is the first example of bracketed comment. SELECT 'ommented out content' AS first -- !query schema struct<> -- !query output org.apache.spark.sql.catalyst.parser.ParseException mismatched input '/' expecting {'(', 'ADD', 'ALTER', 'ANALYZE', 'CACHE', 'CLEAR', 'COMMENT', 'COMMIT', 'CREATE', 'DELETE', 'DESC', 'DESCRIBE', 'DFS', 'DROP', 'EXPLAIN', 'EXPORT', 'FROM', 'GRANT', 'IMPORT', 'INSERT', 'LIST', 'LOAD', 'LOCK', 'MAP', 'MERGE', 'MSCK', 'REDUCE', 'REFRESH', 'REPLACE', 'RESET', 'REVOKE', ' ROLLBACK', 'SELECT', 'SET', 'SHOW', 'START', 'TABLE', 'TRUNCATE', 'UNCACHE', 'UNLOCK', 'UPDATE', 'USE', 'VALUES', 'WITH'}(line 1, pos 0) == SQL == /* This is the first example of bracketed comment. ^^^ SELECT 'ommented out content' AS first -- !query / SELECT 'selected content' AS first -- !query schema struct<> -- !query output org.apache.spark.sql.catalyst.parser.ParseException extraneous input '/' expecting {'(', 'ADD', 'ALTER', 'ANALYZE', 'CACHE', 'CLEAR', 'COMMENT', 'COMMIT', 'CREATE', 'DELETE', 'DESC', 'DESCRIBE', 'DFS', 'DROP', 'EXPLAIN', 'EXPORT', 'FROM', 'GRANT', 'IMPORT', 'INSERT', 'LIST', 'LOAD', 'LOCK', 'MAP', 'MERGE', 'MSCK', 'REDUCE', 'REFRESH', 'REPLACE', 'RESET', 'REVOKE', 'ROLLBACK', 'SELECT', 'SET', 'SHOW', 'START', 'TABLE', 'TRUNCATE', 'UNCACHE', 'UNLOCK', 'UPDATE', 'USE', 'VALUES', 'WITH'}(line 1, pos 0) == SQL == / ^^^ SELECT 'selected content' AS first ``` After this PR, the output will be: ``` -- !query / This is the first example of bracketed comment. SELECT 'ommented out content' AS first; */ SELECT 'selected content' AS first -- !query schema struct<first:string> -- !query output selected content ``` ### Why are the changes needed? Golden files can't display the bracketed comments in imported test cases. ### Does this PR introduce any user-facing change? 'No'. ### How was this patch tested? New UT. Closes #28018 from beliefer/fix-bug-tests-imported-bracketed-comments. Authored-by: beliefer <beliefer@163.com> Signed-off-by: Takeshi Yamamuro <yamamuro@apache.org>	2020-03-27 08:09:17 +09:00
Maxim Gekk	d72ec85741	[SPARK-31238][SQL] Rebase dates to/from Julian calendar in write/read for ORC datasource ### What changes were proposed in this pull request? This PR (SPARK-31238) aims the followings. 1. Modified ORC Vectorized Reader, in particular, OrcColumnVector v1.2 and v2.3. After the changes, it uses `DateTimeUtils. rebaseJulianToGregorianDays()` added by https://github.com/apache/spark/pull/27915 . The method performs rebasing days from the hybrid calendar (Julian + Gregorian) to Proleptic Gregorian calendar. It builds a local date in the original calendar, extracts date fields `year`, `month` and `day` from the local date, and builds another local date in the target calendar. After that, it calculates days from the epoch `1970-01-01` for the resulted local date. 2. Introduced rebasing dates while saving ORC files, in particular, I modified `OrcShimUtils. getDateWritable` v1.2 and v2.3, and returned `DaysWritable` instead of Hive's `DateWritable`. The `DaysWritable` class was added by the PR https://github.com/apache/spark/pull/27890 (and fixed by https://github.com/apache/spark/pull/27962). I moved `DaysWritable` from `sql/hive` to `sql/core` to re-use it in ORC datasource. ### Why are the changes needed? For the backward compatibility with Spark 2.4 and earlier versions. The changes allow users to read dates/timestamps saved by previous version, and get the same result. ### Does this PR introduce any user-facing change? Yes. Before the changes, loading the date `1200-01-01` saved by Spark 2.4.5 returns the following: ```scala scala> spark.read.orc("/Users/maxim/tmp/before_1582/2_4_5_date_orc").show(false) +----------+ \|dt \| +----------+ \|1200-01-08\| +----------+ ``` After the changes ```scala scala> spark.read.orc("/Users/maxim/tmp/before_1582/2_4_5_date_orc").show(false) +----------+ \|dt \| +----------+ \|1200-01-01\| +----------+ ``` ### How was this patch tested? - By running `OrcSourceSuite` and `HiveOrcSourceSuite`. - Add new test `SPARK-31238: compatibility with Spark 2.4 in reading dates` to `OrcSuite` which reads an ORC file saved by Spark 2.4.5 via the commands: ```shell $ export TZ="America/Los_Angeles" ``` ```scala scala> sql("select cast('1200-01-01' as date) dt").write.mode("overwrite").orc("/Users/maxim/tmp/before_1582/2_4_5_date_orc") scala> spark.read.orc("/Users/maxim/tmp/before_1582/2_4_5_date_orc").show(false) +----------+ \|dt \| +----------+ \|1200-01-01\| +----------+ ``` - Add round trip test `SPARK-31238: rebasing dates in write`. The test `SPARK-31238: compatibility with Spark 2.4 in reading dates` confirms rebasing in read. So, we can check rebasing in write. Closes #28016 from MaxGekk/rebase-date-orc. Authored-by: Maxim Gekk <max.gekk@gmail.com> Signed-off-by: Dongjoon Hyun <dongjoon@apache.org>	2020-03-26 13:14:28 -07:00

... 3 4 5 6 7 ...

27074 commits