ODIn/spark-instrumented-optimizer

Author	SHA1	Message	Date
Alexander Ulanov	26867ebc67	[SPARK-11262][ML] Unit test for gradient, loss layers, memory management for multilayer perceptron 1.Implement LossFunction trait and implement squared error and cross entropy loss with it 2.Implement unit test for gradient and loss 3.Implement InPlace trait and in-place layer evaluation 4.Refactor interface for ActivationFunction 5.Update of Layer and LayerModel interfaces 6.Fix random weights assignment 7.Implement memory allocation by MLP model instead of individual layers These features decreased the memory usage and increased flexibility of internal API. Author: Alexander Ulanov <nashb@yandex.ru> Author: avulanov <avulanov@gmail.com> Closes #9229 from avulanov/mlp-refactoring.	2016-03-31 23:48:36 -07:00
Cheng Lian	1b070637fa	[SPARK-14295][SPARK-14274][SQL] Implements buildReader() for LibSVM ## What changes were proposed in this pull request? This PR implements `FileFormat.buildReader()` for the LibSVM data source. Besides that, a new interface method `prepareRead()` is added to `FileFormat`: ```scala def prepareRead( sqlContext: SQLContext, options: Map[String, String], files: Seq[FileStatus]): Map[String, String] = options ``` After migrating from `buildInternalScan()` to `buildReader()`, we lost the opportunity to collect necessary global information, since `buildReader()` works in a per-partition manner. For example, LibSVM needs to infer the total number of features if the `numFeatures` data source option is not set. Any necessary collected global information should be returned using the data source options map. By default, this method just returns the original options untouched. An alternative approach is to absorb `inferSchema()` into `prepareRead()`, since schema inference is also some kind of global information gathering. However, this approach wasn't chosen because schema inference is optional, while `prepareRead()` must be called whenever a `HadoopFsRelation` based data source relation is instantiated. One unaddressed problem is that, when `numFeatures` is absent, now the input data will be scanned twice. The `buildInternalScan()` code path doesn't need to do this because it caches the raw parsed RDD in memory before computing the total number of features. However, with `FileScanRDD`, the raw parsed RDD is created in a different way (e.g. partitioning) from the final RDD. ## How was this patch tested? Tested using existing test suites. Author: Cheng Lian <lian@databricks.com> Closes #12088 from liancheng/spark-14295-libsvm-build-reader.	2016-03-31 23:46:08 -07:00
Zhang, Liye	96941b12f8	[SPARK-14242][CORE][NETWORK] avoid copy in compositeBuffer for frame decoder ## What changes were proposed in this pull request? In this patch, we set the initial `maxNumComponents` to `Integer.MAX_VALUE` instead of the default size ( which is 16) when allocating `compositeBuffer` in `TransportFrameDecoder` because `compositeBuffer` will introduce too many memory copies underlying if `compositeBuffer` is with default `maxNumComponents` when the frame size is large (which result in many transport messages). For details, please refer to [SPARK-14242](https://issues.apache.org/jira/browse/SPARK-14242). ## How was this patch tested? spark unit tests and manual tests. For manual tests, we can reproduce the performance issue with following code: `sc.parallelize(Array(1,2,3),3).mapPartitions(a=>Array(new Array[Double](1024 * 1024 * 50)).iterator).reduce((a,b)=> a).length` It's easy to see the performance gain, both from the running time and CPU usage. Author: Zhang, Liye <liye.zhang@intel.com> Closes #12038 from liyezhang556520/spark-14242.	2016-03-31 20:17:52 -07:00
Davies Liu	f0afafdc5d	[SPARK-14267] [SQL] [PYSPARK] execute multiple Python UDFs within single batch ## What changes were proposed in this pull request? This PR support multiple Python UDFs within single batch, also improve the performance. ```python >>> from pyspark.sql.types import IntegerType >>> sqlContext.registerFunction("double", lambda x: x * 2, IntegerType()) >>> sqlContext.registerFunction("add", lambda x, y: x + y, IntegerType()) >>> sqlContext.sql("SELECT double(add(1, 2)), add(double(2), 1)").explain(True) == Parsed Logical Plan == 'Project [unresolvedalias('double('add(1, 2)), None),unresolvedalias('add('double(2), 1), None)] +- OneRowRelation$ == Analyzed Logical Plan == double(add(1, 2)): int, add(double(2), 1): int Project [double(add(1, 2))#14,add(double(2), 1)#15] +- Project [double(add(1, 2))#14,add(double(2), 1)#15] +- Project [pythonUDF0#16 AS double(add(1, 2))#14,pythonUDF0#18 AS add(double(2), 1)#15] +- EvaluatePython [add(pythonUDF1#17, 1)], [pythonUDF0#18] +- EvaluatePython [double(add(1, 2)),double(2)], [pythonUDF0#16,pythonUDF1#17] +- OneRowRelation$ == Optimized Logical Plan == Project [pythonUDF0#16 AS double(add(1, 2))#14,pythonUDF0#18 AS add(double(2), 1)#15] +- EvaluatePython [add(pythonUDF1#17, 1)], [pythonUDF0#18] +- EvaluatePython [double(add(1, 2)),double(2)], [pythonUDF0#16,pythonUDF1#17] +- OneRowRelation$ == Physical Plan == WholeStageCodegen : +- Project [pythonUDF0#16 AS double(add(1, 2))#14,pythonUDF0#18 AS add(double(2), 1)#15] : +- INPUT +- !BatchPythonEvaluation [add(pythonUDF1#17, 1)], [pythonUDF0#16,pythonUDF1#17,pythonUDF0#18] +- !BatchPythonEvaluation [double(add(1, 2)),double(2)], [pythonUDF0#16,pythonUDF1#17] +- Scan OneRowRelation[] ``` ## How was this patch tested? Added new tests. Using the following script to benchmark 1, 2 and 3 udfs, ``` df = sqlContext.range(1, 1 << 23, 1, 4) double = F.udf(lambda x: x * 2, LongType()) print df.select(double(df.id)).count() print df.select(double(df.id), double(df.id + 1)).count() print df.select(double(df.id), double(df.id + 1), double(df.id + 2)).count() ``` Here is the results: N \| Before \| After \| speed up ---- \|------------ \| -------------\|------ 1 \| 22 s \| 7 s \| 3.1X 2 \| 38 s \| 13 s \| 2.9X 3 \| 58 s \| 16 s \| 3.6X This benchmark ran locally with 4 CPUs. For 3 UDFs, it launched 12 Python before before this patch, 4 process after this patch. After this patch, it will use less memory for multiple UDFs than before (less buffering). Author: Davies Liu <davies@databricks.com> Closes #12057 from davies/multi_udfs.	2016-03-31 16:40:20 -07:00
Sital Kedia	8de201baed	[SPARK-14277][CORE] Upgrade Snappy Java to 1.1.2.4 ## What changes were proposed in this pull request? Upgrade snappy to 1.1.2.4 to improve snappy read/write performance. ## How was this patch tested? Tested by running a job on the cluster and saw 7.5% cpu savings after this change. Author: Sital Kedia <skedia@fb.com> Closes #12096 from sitalkedia/snappyRelease.	2016-03-31 16:06:44 -07:00
Josh Rosen	a7af6cd2ea	[SPARK-14281][TESTS] Fix java8-tests and simplify their build This patch fixes a compilation / build break in Spark's `java8-tests` and refactors their POM to simplify the build. See individual commit messages for more details. Author: Josh Rosen <joshrosen@databricks.com> Closes #12073 from JoshRosen/fix-java8-tests.	2016-03-31 13:52:59 -07:00
sethah	b11887c086	[SPARK-14264][PYSPARK][ML] Add feature importance for GBTs in pyspark ## What changes were proposed in this pull request? Feature importances are exposed in the python API for GBTs. Other changes: * Update the random forest feature importance documentation to not repeat decision tree docstring and instead place a reference to it. ## How was this patch tested? Python doc tests were updated to validate GBT feature importance. Author: sethah <seth.hendrickson16@gmail.com> Closes #12056 from sethah/Pyspark_GBT_feature_importance.	2016-03-31 13:00:10 -07:00
Shixiong Zhu	e785402826	[SPARK-14304][SQL][TESTS] Fix tests that don't create temp files in the `java.io.tmpdir` folder ## What changes were proposed in this pull request? If I press `CTRL-C` when running these tests, the temp files will be left in `sql/core` folder and I need to delete them manually. It's annoying. This PR just moves the temp files to the `java.io.tmpdir` folder and add a name prefix for them. ## How was this patch tested? Existing Jenkins tests Author: Shixiong Zhu <shixiong@databricks.com> Closes #12093 from zsxwing/temp-file.	2016-03-31 12:17:25 -07:00
Michel Lemay	3cfbeb70b1	[SPARK-13710][SHELL][WINDOWS] Fix jline dependency on Windows ## What changes were proposed in this pull request? Exclude jline from curator-recipes since it conflicts with scala 2.11 when running spark-shell. Should not affect scala 2.10 since it is builtin. ## How was this patch tested? Ran spark-shell manually. Author: Michel Lemay <mlemay@gmail.com> Closes #12043 from michellemay/spark-13710-fix-jline-on-windows.	2016-03-31 12:15:32 -07:00
Jo Voordeckers	10508f36ad	[SPARK-11327][MESOS] Dispatcher does not respect all args from the Submit request Supersedes https://github.com/apache/spark/pull/9752 Author: Jo Voordeckers <jo.voordeckers@gmail.com> Author: Iulian Dragos <jaguarul@gmail.com> Closes #10370 from jayv/mesos_cluster_params.	2016-03-31 12:08:10 -07:00
Wenchen Fan	0abee534f0	[SPARK-14069][SQL] Improve SparkStatusTracker to also track executor information ## What changes were proposed in this pull request? Track executor information like host and port, cache size, running tasks. TODO: tests ## How was this patch tested? N/A Author: Wenchen Fan <wenchen@databricks.com> Closes #11888 from cloud-fan/status-tracker.	2016-03-31 12:07:19 -07:00
Michael Gummelt	4d93b653f7	[Docs] Update monitoring.md to accurately describe the history server It looks like the docs were recently updated to reflect the History Server's support for incomplete applications, but they still had wording that suggested only completed applications were viewable. This fixes that. My editor also introduced several whitespace removal changes, that I hope are OK, as text files shouldn't have trailing whitespace. To verify they're purely whitespace changes, add `&w=1` to your browser address. If this isn't acceptable, let me know and I'll update the PR. I also didn't think this required a JIRA. Let me know if I should create one. Not tested Author: Michael Gummelt <mgummelt@mesosphere.io> Closes #12045 from mgummelt/update-history-docs.	2016-03-31 12:06:21 -07:00
jeanlyn	8a333d2da8	[SPARK-14243][CORE] update task metrics when removing blocks ## What changes were proposed in this pull request? This PR try to use `incUpdatedBlockStatuses ` to update the `updatedBlockStatuses ` when removing blocks, making sure `BlockManager` correctly updates `updatedBlockStatuses` ## How was this patch tested? test("updated block statuses") in BlockManagerSuite.scala Author: jeanlyn <jeanlyn92@gmail.com> Closes #12091 from jeanlyn/updateBlock.	2016-03-31 12:04:42 -07:00
gatorsmile	446c45bd87	[SPARK-14182][SQL] Parse DDL Command: Alter View This PR is to provide native parsing support for DDL commands: `Alter View`. Since its AST trees are highly similar to `Alter Table`. Thus, both implementation are integrated into the same one. Based on the Hive DDL document: https://cwiki.apache.org/confluence/display/Hive/LanguageManual+DDL and https://cwiki.apache.org/confluence/display/Hive/PartitionedViews Syntax: ```SQL ALTER VIEW view_name RENAME TO new_view_name ``` - to change the name of a view to a different name Syntax: ```SQL ALTER VIEW view_name SET TBLPROPERTIES ('comment' = new_comment); ``` - to add metadata to a view Syntax: ```SQL ALTER VIEW view_name UNSET TBLPROPERTIES [IF EXISTS] ('comment', 'key') ``` - to remove metadata from a view Syntax: ```SQL ALTER VIEW view_name ADD [IF NOT EXISTS] PARTITION spec1[, PARTITION spec2, ...] ``` - to add the partitioning metadata for a view. - the syntax of partition spec in `ALTER VIEW` is identical to `ALTER TABLE`, EXCEPT that it is ILLEGAL to specify a `LOCATION` clause. Syntax: ```SQL ALTER VIEW view_name DROP [IF EXISTS] PARTITION spec1[, PARTITION spec2, ...] ``` - to drop the related partition metadata for a view. Added the related test cases to `DDLCommandSuite` Author: gatorsmile <gatorsmile@gmail.com> Author: xiaoli <lixiao1983@gmail.com> Author: Xiao Li <xiaoli@Xiaos-MacBook-Pro.local> Closes #11987 from gatorsmile/parseAlterView.	2016-03-31 12:04:03 -07:00
Nishkam Ravi	ac1b8b302a	[SPARK-13796] Redirect error message to logWarning ## What changes were proposed in this pull request? Redirect error message to logWarning ## How was this patch tested? Unit tests, manual tests JoshRosen Author: Nishkam Ravi <nishkamravi@gmail.com> Closes #12052 from nishkamravi2/master_warning.	2016-03-31 12:03:05 -07:00
Sameer Agarwal	3586929320	[SPARK-14278][SQL] Initialize columnar batch with proper memory mode ## What changes were proposed in this pull request? Fixes a minor bug in the record reader constructor that was possibly introduced during refactoring. ## How was this patch tested? N/A Author: Sameer Agarwal <sameer@databricks.com> Closes #12070 from sameeragarwal/vectorized-rr.	2016-03-31 11:56:28 -07:00
Sameer Agarwal	8d6207206c	[SPARK-14263][SQL] Benchmark Vectorized HashMap for GroupBy Aggregates ## What changes were proposed in this pull request? This PR proposes a new data-structure based on a vectorized hashmap that can be potentially _codegened_ in `TungstenAggregate` to speed up aggregates with group by. Micro-benchmarks show a 10x improvement over the current `BytesToBytes` aggregation map. ## How was this patch tested? Intel(R) Core(TM) i7-4960HQ CPU 2.60GHz BytesToBytesMap: Best/Avg Time(ms) Rate(M/s) Per Row(ns) Relative ------------------------------------------------------------------------------------------- hash 108 / 119 96.9 10.3 1.0X fast hash 63 / 70 166.2 6.0 1.7X arrayEqual 70 / 73 150.8 6.6 1.6X Java HashMap (Long) 141 / 200 74.3 13.5 0.8X Java HashMap (two ints) 145 / 185 72.3 13.8 0.7X Java HashMap (UnsafeRow) 499 / 524 21.0 47.6 0.2X BytesToBytesMap (off Heap) 483 / 548 21.7 46.0 0.2X BytesToBytesMap (on Heap) 485 / 562 21.6 46.2 0.2X Vectorized Hashmap 54 / 60 193.7 5.2 2.0X Author: Sameer Agarwal <sameer@databricks.com> Closes #12055 from sameeragarwal/vectorized-hashmap.	2016-03-31 11:53:13 -07:00
Xusen Yin	8b207f3b6a	[SPARK-11892][ML] Model export/import for spark.ml: OneVsRest # What changes were proposed in this pull request? https://issues.apache.org/jira/browse/SPARK-11892 Add save/load for spark ml.OneVsRest and its model. Also add OneVsRest and OneVsRestModel in MetaAlgorithmReadWrite. # How was this patch tested? Test with Scala unit test. Author: Xusen Yin <yinxusen@gmail.com> Closes #9934 from yinxusen/SPARK-11892.	2016-03-31 11:17:32 -07:00
Yuhao Yang	a0a1991580	[SPARK-13782][ML] Model export/import for spark.ml: BisectingKMeans ## What changes were proposed in this pull request? jira: https://issues.apache.org/jira/browse/SPARK-13782 Model export/import for BisectingKMeans in spark.ml and mllib ## How was this patch tested? unit tests Author: Yuhao Yang <hhbyyh@gmail.com> Closes #11933 from hhbyyh/bisectingsave.	2016-03-31 11:12:40 -07:00
jerryshao	3b3cc76004	[SPARK-14062][YARN] Fix log4j and upload metrics.properties automatically with distributed cache ## What changes were proposed in this pull request? 1. Currently log4j which uses distributed cache only adds to AM's classpath, not executor's, this is introduced in #9118, which breaks the original meaning of that PR, so here add log4j file to the classpath of both AM and executors. 2. Automatically upload metrics.properties to distributed cache, so that it could be used by remote driver and executors implicitly. ## How was this patch tested? Unit test and integration test is done. Author: jerryshao <sshao@hortonworks.com> Closes #11885 from jerryshao/SPARK-14062.	2016-03-31 10:27:33 -07:00
Dongjoon Hyun	208fff3ac8	[SPARK-14164][MLLIB] Improve input layer validation of MultilayerPerceptronClassifier ## What changes were proposed in this pull request? This issue improves an input layer validation and adds related testcases to MultilayerPerceptronClassifier. ```scala - // TODO: how to check ALSO that all elements are greater than 0? - ParamValidators.arrayLengthGt(1) + (t: Array[Int]) => t.forall(ParamValidators.gt(0)) && t.length > 1 ``` ## How was this patch tested? Pass the Jenkins tests including the new testcases. Author: Dongjoon Hyun <dongjoon@apache.org> Closes #11964 from dongjoon-hyun/SPARK-14164.	2016-03-31 09:39:15 -07:00
Herman van Hovell	a9b93e0739	[SPARK-14211][SQL] Remove ANTLR3 based parser ### What changes were proposed in this pull request? This PR removes the ANTLR3 based parser, and moves the new ANTLR4 based parser into the `org.apache.spark.sql.catalyst.parser package`. ### How was this patch tested? Existing unit tests. cc rxin andrewor14 yhuai Author: Herman van Hovell <hvanhovell@questtec.nl> Closes #12071 from hvanhovell/SPARK-14211.	2016-03-31 09:25:09 -07:00
Cheng Lian	26445c2e47	[SPARK-14206][SQL] buildReader() implementation for CSV ## What changes were proposed in this pull request? Major changes: 1. Implement `FileFormat.buildReader()` for the CSV data source. 1. Add an extra argument to `FileFormat.buildReader()`, `physicalSchema`, which is basically the result of `FileFormat.inferSchema` or user specified schema. This argument is necessary because the CSV data source needs to know all the columns of the underlying files to read the file. ## How was this patch tested? Existing tests should do the work. Author: Cheng Lian <lian@databricks.com> Closes #12002 from liancheng/spark-14206-csv-build-reader.	2016-03-30 18:21:06 -07:00
Travis Crawford	da54abfd87	[SPARK-14081][SQL] - Preserve DataFrame column types when filling nulls. ## What changes were proposed in this pull request? This change resolves an issue where `DataFrameNaFunctions.fill` changes a `FloatType` column to a `DoubleType`. We also clarify the contract that replacement values will be cast to the column data type, which may change the replacement value when casting to a lower precision type. ## How was this patch tested? This patch has associated unit tests. Author: Travis Crawford <travis@medium.com> Closes #11967 from traviscrawford/SPARK-14081-dataframena.	2016-03-30 16:59:52 -07:00
Dongjoon Hyun	258a243419	[SPARK-14282][SQL] CodeFormatter should handle oneline comment with /* / properly ## What changes were proposed in this pull request? This PR improves `CodeFormatter` to fix the following malformed indentations. ```java / 019 / public java.lang.Object apply(java.lang.Object _i) { / 020 / InternalRow i = (InternalRow) _i; / 021 / / createexternalrow(if (isnull(input[0, double])) null else input[0, double], if (isnull(input[1, int])) null else input[1, int], ... / / 022 / boolean isNull = false; / 023 / final Object[] values = new Object[2]; / 024 / / if (isnull(input[0, double])) null else input[0, double] / / 025 / / isnull(input[0, double]) / ... / 053 / if (!false && false) { / 054 / / null / / 055 / final int value9 = -1; / 056 / isNull6 = true; / 057 / value6 = value9; / 058 / } else { ... / 077 / return mutableRow; / 078 / } / 079 / } / 080 / ``` After this PR, the code will be formatted like the following. ```java / 019 / public java.lang.Object apply(java.lang.Object _i) { / 020 / InternalRow i = (InternalRow) _i; / 021 / / createexternalrow(if (isnull(input[0, double])) null else input[0, double], if (isnull(input[1, int])) null else input[1, int], ... / / 022 / boolean isNull = false; / 023 / final Object[] values = new Object[2]; / 024 / / if (isnull(input[0, double])) null else input[0, double] / / 025 / / isnull(input[0, double]) / ... / 053 / if (!false && false) { / 054 / / null / / 055 / final int value9 = -1; / 056 / isNull6 = true; / 057 / value6 = value9; / 058 / } else { ... / 077 / return mutableRow; / 078 / } / 079 / } / 080 / ``` Also, this issue fixes the following too. (Similar with [SPARK-14185](https://issues.apache.org/jira/browse/SPARK-14185)) ```java 16/03/30 12:39:24 DEBUG WholeStageCodegen: / 001 / public Object generate(Object[] references) { / 002 / return new GeneratedIterator(references); / 003 / } ``` ```java 16/03/30 12:46:32 DEBUG WholeStageCodegen: / 001 / public Object generate(Object[] references) { / 002 / return new GeneratedIterator(references); / 003 */ } ``` ## How was this patch tested? Pass the Jenkins tests (including new CodeFormatterSuite testcases.) Author: Dongjoon Hyun <dongjoon@apache.org> Closes #12072 from dongjoon-hyun/SPARK-14282.	2016-03-30 16:15:37 -07:00
Takeshi YAMAMURO	dadf0138b3	[SPARK-14259][SQL] Add a FileSourceStrategy option for limiting #files in a partition ## What changes were proposed in this pull request? This pr is to add a config to control the maximum number of files as even small files have a non-trivial fixed cost. The current packing can put a lot of small files together which cases straggler tasks. ## How was this patch tested? I added tests to check if many files get split into partitions in FileSourceStrategySuite. Author: Takeshi YAMAMURO <linguin.m.s@gmail.com> Closes #12068 from maropu/SPARK-14259.	2016-03-30 16:02:48 -07:00
Yuhao Yang	ca458618d8	[SPARK-11507][MLLIB] add compact in Matrices fromBreeze jira: https://issues.apache.org/jira/browse/SPARK-11507 "In certain situations when adding two block matrices, I get an error regarding colPtr and the operation fails. External issue URL includes full error and code for reproducing the problem." root cause: colPtr.last does NOT always equal to values.length in breeze SCSMatrix, which fails the require in SparseMatrix. easy step to repro: ``` val m1: BM[Double] = new CSCMatrix[Double] (Array (1.0, 1, 1), 3, 3, Array (0, 1, 2, 3), Array (0, 1, 2) ) val m2: BM[Double] = new CSCMatrix[Double] (Array (1.0, 2, 2, 4), 3, 3, Array (0, 0, 2, 4), Array (1, 2, 1, 2) ) val sum = m1 + m2 Matrices.fromBreeze(sum) ``` Solution: By checking the code in [CSCMatrix](`28000a7b90/math/src/main/scala/breeze/linalg/CSCMatrix.scala`), CSCMatrix in breeze can have extra zeros in the end of data array. Invoking compact will make sure it aligns with the require of SparseMatrix. This should add limited overhead as the actual compact operation is only performed when necessary. Author: Yuhao Yang <hhbyyh@gmail.com> Closes #9520 from hhbyyh/matricesFromBreeze.	2016-03-30 15:58:19 -07:00
Yanbo Liang	f301df37cb	[SPARK-14152][ML][PYSPARK] MultilayerPerceptronClassifier supports save/load for Python API ## What changes were proposed in this pull request? ```MultilayerPerceptronClassifier``` supports save/load for Python API. ## How was this patch tested? doctest. cc mengxr jkbradley yinxusen Author: Yanbo Liang <ybliang8@gmail.com> Closes #11952 from yanboliang/spark-14152.	2016-03-30 15:47:01 -07:00
Yanbo Liang	5dc948e812	[MINOR][ML] Fix the wrong param name of LDA topicDistributionCol ## What changes were proposed in this pull request? Fix the wrong param name of LDA ```topicDistributionCol```. ## How was this patch tested? No tests. cc jkbradley Author: Yanbo Liang <ybliang8@gmail.com> Closes #12065 from yanboliang/lda-topicDistributionCol.	2016-03-30 14:57:38 -07:00
Xusen Yin	529d6ce8f9	[SPARK-14181] TrainValidationSplit should have HasSeed https://issues.apache.org/jira/browse/SPARK-14181 TrainValidationSplit should have HasSeed for the random split of RDD. I also changed the random split from the RDD function to the DataFrame function. Author: Xusen Yin <yinxusen@gmail.com> Closes #11985 from yinxusen/SPARK-14181.	2016-03-30 14:32:29 -07:00
Marcelo Vanzin	bdabfd43f6	[SPARK-13955][YARN] Also look for Spark jars in the build directory. Move the logic to find Spark jars to CommandBuilderUtils and make it available for YARN code, so that it's possible to easily launch Spark on YARN from a build directory. Tested by running SparkPi from the build directory on YARN. Author: Marcelo Vanzin <vanzin@cloudera.com> Closes #11970 from vanzin/SPARK-13955.	2016-03-30 13:59:10 -07:00
Wenchen Fan	d46c71b39d	[SPARK-14268][SQL] rename toRowExpressions and fromRowExpression to serializer and deserializer in ExpressionEncoder ## What changes were proposed in this pull request? In `ExpressionEncoder`, we use `constructorFor` to build `fromRowExpression` as the `deserializer` in `ObjectOperator`. It's kind of confusing, we should make the name consistent. ## How was this patch tested? existing tests. Author: Wenchen Fan <wenchen@databricks.com> Closes #12058 from cloud-fan/rename.	2016-03-30 11:03:15 -07:00
Wenchen Fan	816f359cf0	[SPARK-14114][SQL] implement buildReader for text data source ## What changes were proposed in this pull request? This PR implements buildReader for text data source and enable it in the new data source code path. ## How was this patch tested? Existing tests. Author: Wenchen Fan <wenchen@databricks.com> Closes #11934 from cloud-fan/text.	2016-03-30 17:32:53 +08:00
Shixiong Zhu	7320f9bd19	[SPARK-14254][CORE] Add logs to help investigate the network performance ## What changes were proposed in this pull request? It would be very helpful for network performance investigation if we log the time spent on connecting and resolving host. ## How was this patch tested? Jenkins unit tests. Author: Shixiong Zhu <shixiong@databricks.com> Closes #12046 from zsxwing/connection-time.	2016-03-29 21:14:48 -07:00
gatorsmile	b66b97cd04	[SPARK-14124][SQL] Implement Database-related DDL Commands #### What changes were proposed in this pull request? This PR is to implement the following four Database-related DDL commands: - `CREATE DATABASE\|SCHEMA [IF NOT EXISTS] database_name` - `DROP DATABASE [IF EXISTS] database_name [RESTRICT\|CASCADE]` - `DESCRIBE DATABASE [EXTENDED] db_name` - `ALTER (DATABASE\|SCHEMA) database_name SET DBPROPERTIES (property_name=property_value, ...)` Another PR will be submitted to handle the unsupported commands. In the Database-related DDL commands, we will issue an error exception for `ALTER (DATABASE\|SCHEMA) database_name SET OWNER [USER\|ROLE] user_or_role`. cc yhuai andrewor14 rxin Could you review the changes? Is it in the right direction? Thanks! #### How was this patch tested? Added a few test cases in `command/DDLSuite.scala` for testing DDL command execution in `SQLContext`. Since `HiveContext` also shares the same implementation, the existing test cases in `\hive` also verifies the correctness of these commands. Author: gatorsmile <gatorsmile@gmail.com> Author: xiaoli <lixiao1983@gmail.com> Author: Xiao Li <xiaoli@Xiaos-MacBook-Pro.local> Closes #12009 from gatorsmile/dbDDL.	2016-03-29 17:39:52 -07:00
tedyu	e1f6845391	[SPARK-12181] Check Cached unaligned-access capability before using Unsafe ## What changes were proposed in this pull request? For MemoryMode.OFF_HEAP, Unsafe.getInt etc. are used with no restriction. However, the Oracle implementation uses these methods only if the class variable unaligned (commented as "Cached unaligned-access capability") is true, which seems to be calculated whether the architecture is i386, x86, amd64, or x86_64. I think we should perform similar check for the use of Unsafe. Reference: https://github.com/netty/netty/blob/4.1/common/src/main/java/io/netty/util/internal/PlatformDependent0.java#L112 ## How was this patch tested? Unit test suite Author: tedyu <yuzhihong@gmail.com> Closes #11943 from tedyu/master.	2016-03-29 17:16:53 -07:00
Sameer Agarwal	366cac6fb0	[SPARK-14225][SQL] Cap the length of toCommentSafeString at 128 chars ## What changes were proposed in this pull request? Builds on https://github.com/apache/spark/pull/12022 and (a) appends "..." to truncated comment strings and (b) fixes indentation in lines after the commented strings if they happen to have a `(`, `{`, `)` or `}` ## How was this patch tested? Manually examined the generated code. Author: Sameer Agarwal <sameer@databricks.com> Closes #12044 from sameeragarwal/comment.	2016-03-29 16:46:45 -07:00
Davies Liu	a7a93a116d	[SPARK-14215] [SQL] [PYSPARK] Support chained Python UDFs ## What changes were proposed in this pull request? This PR brings the support for chained Python UDFs, for example ```sql select udf1(udf2(a)) select udf1(udf2(a) + 3) select udf1(udf2(a) + udf3(b)) ``` Also directly chained unary Python UDFs are put in single batch of Python UDFs, others may require multiple batches. For example, ```python >>> sqlContext.sql("select double(double(1))").explain() == Physical Plan == WholeStageCodegen : +- Project [pythonUDF#10 AS double(double(1))#9] : +- INPUT +- !BatchPythonEvaluation double(double(1)), [pythonUDF#10] +- Scan OneRowRelation[] >>> sqlContext.sql("select double(double(1) + double(2))").explain() == Physical Plan == WholeStageCodegen : +- Project [pythonUDF#19 AS double((double(1) + double(2)))#16] : +- INPUT +- !BatchPythonEvaluation double((pythonUDF#17 + pythonUDF#18)), [pythonUDF#17,pythonUDF#18,pythonUDF#19] +- !BatchPythonEvaluation double(2), [pythonUDF#17,pythonUDF#18] +- !BatchPythonEvaluation double(1), [pythonUDF#17] +- Scan OneRowRelation[] ``` TODO: will support multiple unrelated Python UDFs in one batch (another PR). ## How was this patch tested? Added new unit tests for chained UDFs. Author: Davies Liu <davies@databricks.com> Closes #12014 from davies/py_udfs.	2016-03-29 15:06:29 -07:00
Eric Liang	e58c4cb3c5	[SPARK-14227][SQL] Add method for printing out generated code for debugging ## What changes were proposed in this pull request? This adds `debugCodegen` to the debug package for query execution. ## How was this patch tested? Unit and manual testing. Output example: ``` scala> import org.apache.spark.sql.execution.debug._ import org.apache.spark.sql.execution.debug._ scala> sqlContext.range(100).groupBy("id").count().orderBy("id").debugCodegen() Found 3 WholeStageCodegen subtrees. == Subtree 1 / 3 == WholeStageCodegen : +- TungstenAggregate(key=[id#0L], functions=[(count(1),mode=Partial,isDistinct=false)], output=[id#0L,count#9L]) : +- Range 0, 1, 1, 100, [id#0L] Generated code: /* 001 / public Object generate(Object[] references) { / 002 / return new GeneratedIterator(references); / 003 / } / 004 / / 005 / /* Codegened pipeline for: /* 006 / TungstenAggregate(key=[id#0L], functions=[(count(1),mode=Partial,isDistinct=false)], output=[id#0L,count#9L]) /* 007 / +- Range 0, 1, 1, 100, [id#0L] / 008 / / /* 009 / final class GeneratedIterator extends org.apache.spark.sql.execution.BufferedRowIterator { / 010 / private Object[] references; / 011 / private boolean agg_initAgg; / 012 / private org.apache.spark.sql.execution.aggregate.TungstenAggregate agg_plan; / 013 / private org.apache.spark.sql.execution.UnsafeFixedWidthAggregationMap agg_hashMap; / 014 / private org.apache.spark.sql.execution.UnsafeKVExternalSorter agg_sorter; / 015 / private org.apache.spark.unsafe.KVIterator agg_mapIter; / 016 / private org.apache.spark.sql.execution.metric.LongSQLMetric range_numOutputRows; / 017 / private org.apache.spark.sql.execution.metric.LongSQLMetricValue range_metricValue; / 018 / private boolean range_initRange; / 019 / private long range_partitionEnd; / 020 / private long range_number; / 021 / private boolean range_overflow; / 022 / private scala.collection.Iterator range_input; / 023 / private UnsafeRow range_result; / 024 / private org.apache.spark.sql.catalyst.expressions.codegen.BufferHolder range_holder; / 025 / private org.apache.spark.sql.catalyst.expressions.codegen.UnsafeRowWriter range_rowWriter; / 026 / private UnsafeRow agg_result; / 027 / private org.apache.spark.sql.catalyst.expressions.codegen.BufferHolder agg_holder; / 028 / private org.apache.spark.sql.catalyst.expressions.codegen.UnsafeRowWriter agg_rowWriter; / 029 / private org.apache.spark.sql.catalyst.expressions.codegen.UnsafeRowJoiner agg_unsafeRowJoiner; / 030 / private org.apache.spark.sql.execution.metric.LongSQLMetric wholestagecodegen_numOutputRows; / 031 / private org.apache.spark.sql.execution.metric.LongSQLMetricValue wholestagecodegen_metricValue; / 032 / / 033 / public GeneratedIterator(Object[] references) { / 034 / this.references = references; / 035 / } / 036 / / 037 / public void init(scala.collection.Iterator inputs[]) { / 038 / agg_initAgg = false; / 039 / this.agg_plan = (org.apache.spark.sql.execution.aggregate.TungstenAggregate) references[0]; / 040 / agg_hashMap = agg_plan.createHashMap(); / 041 / / 042 / this.range_numOutputRows = (org.apache.spark.sql.execution.metric.LongSQLMetric) references[1]; / 043 / range_metricValue = (org.apache.spark.sql.execution.metric.LongSQLMetricValue) range_numOutputRows.localValue(); / 044 / range_initRange = false; / 045 / range_partitionEnd = 0L; / 046 / range_number = 0L; / 047 / range_overflow = false; / 048 / range_input = inputs[0]; / 049 / range_result = new UnsafeRow(1); / 050 / this.range_holder = new org.apache.spark.sql.catalyst.expressions.codegen.BufferHolder(range_result, 0); / 051 / this.range_rowWriter = new org.apache.spark.sql.catalyst.expressions.codegen.UnsafeRowWriter(range_holder, 1); / 052 / agg_result = new UnsafeRow(1); / 053 / this.agg_holder = new org.apache.spark.sql.catalyst.expressions.codegen.BufferHolder(agg_result, 0); / 054 / this.agg_rowWriter = new org.apache.spark.sql.catalyst.expressions.codegen.UnsafeRowWriter(agg_holder, 1); / 055 / agg_unsafeRowJoiner = agg_plan.createUnsafeJoiner(); / 056 / this.wholestagecodegen_numOutputRows = (org.apache.spark.sql.execution.metric.LongSQLMetric) references[2]; / 057 / wholestagecodegen_metricValue = (org.apache.spark.sql.execution.metric.LongSQLMetricValue) wholestagecodegen_numOutputRows.localValue(); / 058 / } / 059 / / 060 / private void agg_doAggregateWithKeys() throws java.io.IOException { / 061 / /** PRODUCE: Range 0, 1, 1, 100, [id#0L] / / 062 / / 063 / // initialize Range / 064 / if (!range_initRange) { / 065 / range_initRange = true; / 066 / if (range_input.hasNext()) { / 067 / initRange(((InternalRow) range_input.next()).getInt(0)); / 068 / } else { / 069 / return; / 070 / } / 071 / } / 072 / / 073 / while (!range_overflow && range_number < range_partitionEnd) { / 074 / long range_value = range_number; / 075 / range_number += 1L; / 076 / if (range_number < range_value ^ 1L < 0) { / 077 / range_overflow = true; / 078 / } / 079 / / 080 / /** CONSUME: TungstenAggregate(key=[id#0L], functions=[(count(1),mode=Partial,isDistinct=false)], output=[id#0L,count#9L]) / / 081 / / 082 / // generate grouping key / 083 / agg_rowWriter.write(0, range_value); / 084 / / hash(input[0, bigint], 42) / / 085 / int agg_value1 = 42; / 086 / / 087 / agg_value1 = org.apache.spark.unsafe.hash.Murmur3_x86_32.hashLong(range_value, agg_value1); / 088 / UnsafeRow agg_aggBuffer = null; / 089 / if (true) { / 090 / // try to get the buffer from hash map / 091 / agg_aggBuffer = agg_hashMap.getAggregationBufferFromUnsafeRow(agg_result, agg_value1); / 092 / } / 093 / if (agg_aggBuffer == null) { / 094 / if (agg_sorter == null) { / 095 / agg_sorter = agg_hashMap.destructAndCreateExternalSorter(); / 096 / } else { / 097 / agg_sorter.merge(agg_hashMap.destructAndCreateExternalSorter()); / 098 / } / 099 / / 100 / // the hash map had be spilled, it should have enough memory now, / 101 / // try to allocate buffer again. / 102 / agg_aggBuffer = agg_hashMap.getAggregationBufferFromUnsafeRow(agg_result, agg_value1); / 103 / if (agg_aggBuffer == null) { / 104 / // failed to allocate the first page / 105 / throw new OutOfMemoryError("No enough memory for aggregation"); / 106 / } / 107 / } / 108 / / 109 / // evaluate aggregate function / 110 / / (input[0, bigint] + 1) / / 111 / / input[0, bigint] / / 112 / long agg_value4 = agg_aggBuffer.getLong(0); / 113 / / 114 / long agg_value3 = -1L; / 115 / agg_value3 = agg_value4 + 1L; / 116 / // update aggregate buffer / 117 / agg_aggBuffer.setLong(0, agg_value3); / 118 / / 119 / if (shouldStop()) return; / 120 / } / 121 / / 122 / agg_mapIter = agg_plan.finishAggregate(agg_hashMap, agg_sorter); / 123 / } / 124 / / 125 / private void initRange(int idx) { / 126 / java.math.BigInteger index = java.math.BigInteger.valueOf(idx); / 127 / java.math.BigInteger numSlice = java.math.BigInteger.valueOf(1L); / 128 / java.math.BigInteger numElement = java.math.BigInteger.valueOf(100L); / 129 / java.math.BigInteger step = java.math.BigInteger.valueOf(1L); / 130 / java.math.BigInteger start = java.math.BigInteger.valueOf(0L); / 131 / / 132 / java.math.BigInteger st = index.multiply(numElement).divide(numSlice).multiply(step).add(start); / 133 / if (st.compareTo(java.math.BigInteger.valueOf(Long.MAX_VALUE)) > 0) { / 134 / range_number = Long.MAX_VALUE; / 135 / } else if (st.compareTo(java.math.BigInteger.valueOf(Long.MIN_VALUE)) < 0) { / 136 / range_number = Long.MIN_VALUE; / 137 / } else { / 138 / range_number = st.longValue(); / 139 / } / 140 / / 141 / java.math.BigInteger end = index.add(java.math.BigInteger.ONE).multiply(numElement).divide(numSlice) / 142 / .multiply(step).add(start); / 143 / if (end.compareTo(java.math.BigInteger.valueOf(Long.MAX_VALUE)) > 0) { / 144 / range_partitionEnd = Long.MAX_VALUE; / 145 / } else if (end.compareTo(java.math.BigInteger.valueOf(Long.MIN_VALUE)) < 0) { / 146 / range_partitionEnd = Long.MIN_VALUE; / 147 / } else { / 148 / range_partitionEnd = end.longValue(); / 149 / } / 150 / / 151 / range_metricValue.add((range_partitionEnd - range_number) / 1L); / 152 / } / 153 / / 154 / protected void processNext() throws java.io.IOException { / 155 / /** PRODUCE: TungstenAggregate(key=[id#0L], functions=[(count(1),mode=Partial,isDistinct=false)], output=[id#0L,count#9L]) / / 156 / / 157 / if (!agg_initAgg) { / 158 / agg_initAgg = true; / 159 / agg_doAggregateWithKeys(); / 160 / } / 161 / / 162 / // output the result / 163 / while (agg_mapIter.next()) { / 164 / wholestagecodegen_metricValue.add(1); / 165 / UnsafeRow agg_aggKey = (UnsafeRow) agg_mapIter.getKey(); / 166 / UnsafeRow agg_aggBuffer1 = (UnsafeRow) agg_mapIter.getValue(); / 167 / / 168 / UnsafeRow agg_resultRow = agg_unsafeRowJoiner.join(agg_aggKey, agg_aggBuffer1); / 169 / / 170 / /** CONSUME: WholeStageCodegen / / 171 / / 172 / append(agg_resultRow); / 173 / / 174 / if (shouldStop()) return; / 175 / } / 176 / / 177 / agg_mapIter.close(); / 178 / if (agg_sorter == null) { / 179 / agg_hashMap.free(); / 180 / } / 181 / } / 182 / } == Subtree 2 / 3 == WholeStageCodegen : +- Sort [id#0L ASC], true, 0 : +- INPUT +- Exchange rangepartitioning(id#0L ASC, 200), None +- WholeStageCodegen : +- TungstenAggregate(key=[id#0L], functions=[(count(1),mode=Final,isDistinct=false)], output=[id#0L,count#4L]) : +- INPUT +- Exchange hashpartitioning(id#0L, 200), None +- WholeStageCodegen : +- TungstenAggregate(key=[id#0L], functions=[(count(1),mode=Partial,isDistinct=false)], output=[id#0L,count#9L]) : +- Range 0, 1, 1, 100, [id#0L] Generated code: / 001 / public Object generate(Object[] references) { / 002 / return new GeneratedIterator(references); / 003 / } / 004 / / 005 / /* Codegened pipeline for: /* 006 / Sort [id#0L ASC], true, 0 /* 007 / +- INPUT / 008 / / /* 009 / final class GeneratedIterator extends org.apache.spark.sql.execution.BufferedRowIterator { / 010 / private Object[] references; / 011 / private boolean sort_needToSort; / 012 / private org.apache.spark.sql.execution.Sort sort_plan; / 013 / private org.apache.spark.sql.execution.UnsafeExternalRowSorter sort_sorter; / 014 / private org.apache.spark.executor.TaskMetrics sort_metrics; / 015 / private scala.collection.Iterator<UnsafeRow> sort_sortedIter; / 016 / private scala.collection.Iterator inputadapter_input; / 017 / private org.apache.spark.sql.execution.metric.LongSQLMetric sort_dataSize; / 018 / private org.apache.spark.sql.execution.metric.LongSQLMetricValue sort_metricValue; / 019 / private org.apache.spark.sql.execution.metric.LongSQLMetric sort_spillSize; / 020 / private org.apache.spark.sql.execution.metric.LongSQLMetricValue sort_metricValue1; / 021 / / 022 / public GeneratedIterator(Object[] references) { / 023 / this.references = references; / 024 / } / 025 / / 026 / public void init(scala.collection.Iterator inputs[]) { / 027 / sort_needToSort = true; / 028 / this.sort_plan = (org.apache.spark.sql.execution.Sort) references[0]; / 029 / sort_sorter = sort_plan.createSorter(); / 030 / sort_metrics = org.apache.spark.TaskContext.get().taskMetrics(); / 031 / / 032 / inputadapter_input = inputs[0]; / 033 / this.sort_dataSize = (org.apache.spark.sql.execution.metric.LongSQLMetric) references[1]; / 034 / sort_metricValue = (org.apache.spark.sql.execution.metric.LongSQLMetricValue) sort_dataSize.localValue(); / 035 / this.sort_spillSize = (org.apache.spark.sql.execution.metric.LongSQLMetric) references[2]; / 036 / sort_metricValue1 = (org.apache.spark.sql.execution.metric.LongSQLMetricValue) sort_spillSize.localValue(); / 037 / } / 038 / / 039 / private void sort_addToSorter() throws java.io.IOException { / 040 / /** PRODUCE: INPUT / / 041 / / 042 / while (inputadapter_input.hasNext()) { / 043 / InternalRow inputadapter_row = (InternalRow) inputadapter_input.next(); / 044 / /** CONSUME: Sort [id#0L ASC], true, 0 / / 045 / / 046 / sort_sorter.insertRow((UnsafeRow)inputadapter_row); / 047 / if (shouldStop()) return; / 048 / } / 049 / / 050 / } / 051 / / 052 / protected void processNext() throws java.io.IOException { / 053 / /** PRODUCE: Sort [id#0L ASC], true, 0 / / 054 / if (sort_needToSort) { / 055 / sort_addToSorter(); / 056 / Long sort_spillSizeBefore = sort_metrics.memoryBytesSpilled(); / 057 / sort_sortedIter = sort_sorter.sort(); / 058 / sort_metricValue.add(sort_sorter.getPeakMemoryUsage()); / 059 / sort_metricValue1.add(sort_metrics.memoryBytesSpilled() - sort_spillSizeBefore); / 060 / sort_metrics.incPeakExecutionMemory(sort_sorter.getPeakMemoryUsage()); / 061 / sort_needToSort = false; / 062 / } / 063 / / 064 / while (sort_sortedIter.hasNext()) { / 065 / UnsafeRow sort_outputRow = (UnsafeRow)sort_sortedIter.next(); / 066 / / 067 / /** CONSUME: WholeStageCodegen / / 068 / / 069 / append(sort_outputRow); / 070 / / 071 / if (shouldStop()) return; / 072 / } / 073 / } / 074 / } == Subtree 3 / 3 == WholeStageCodegen : +- TungstenAggregate(key=[id#0L], functions=[(count(1),mode=Final,isDistinct=false)], output=[id#0L,count#4L]) : +- INPUT +- Exchange hashpartitioning(id#0L, 200), None +- WholeStageCodegen : +- TungstenAggregate(key=[id#0L], functions=[(count(1),mode=Partial,isDistinct=false)], output=[id#0L,count#9L]) : +- Range 0, 1, 1, 100, [id#0L] Generated code: / 001 / public Object generate(Object[] references) { / 002 / return new GeneratedIterator(references); / 003 / } / 004 / / 005 / /* Codegened pipeline for: /* 006 / TungstenAggregate(key=[id#0L], functions=[(count(1),mode=Final,isDistinct=false)], output=[id#0L,count#4L]) /* 007 / +- INPUT / 008 / / /* 009 / final class GeneratedIterator extends org.apache.spark.sql.execution.BufferedRowIterator { / 010 / private Object[] references; / 011 / private boolean agg_initAgg; / 012 / private org.apache.spark.sql.execution.aggregate.TungstenAggregate agg_plan; / 013 / private org.apache.spark.sql.execution.UnsafeFixedWidthAggregationMap agg_hashMap; / 014 / private org.apache.spark.sql.execution.UnsafeKVExternalSorter agg_sorter; / 015 / private org.apache.spark.unsafe.KVIterator agg_mapIter; / 016 / private scala.collection.Iterator inputadapter_input; / 017 / private UnsafeRow agg_result; / 018 / private org.apache.spark.sql.catalyst.expressions.codegen.BufferHolder agg_holder; / 019 / private org.apache.spark.sql.catalyst.expressions.codegen.UnsafeRowWriter agg_rowWriter; / 020 / private UnsafeRow agg_result1; / 021 / private org.apache.spark.sql.catalyst.expressions.codegen.BufferHolder agg_holder1; / 022 / private org.apache.spark.sql.catalyst.expressions.codegen.UnsafeRowWriter agg_rowWriter1; / 023 / private org.apache.spark.sql.execution.metric.LongSQLMetric wholestagecodegen_numOutputRows; / 024 / private org.apache.spark.sql.execution.metric.LongSQLMetricValue wholestagecodegen_metricValue; / 025 / / 026 / public GeneratedIterator(Object[] references) { / 027 / this.references = references; / 028 / } / 029 / / 030 / public void init(scala.collection.Iterator inputs[]) { / 031 / agg_initAgg = false; / 032 / this.agg_plan = (org.apache.spark.sql.execution.aggregate.TungstenAggregate) references[0]; / 033 / agg_hashMap = agg_plan.createHashMap(); / 034 / / 035 / inputadapter_input = inputs[0]; / 036 / agg_result = new UnsafeRow(1); / 037 / this.agg_holder = new org.apache.spark.sql.catalyst.expressions.codegen.BufferHolder(agg_result, 0); / 038 / this.agg_rowWriter = new org.apache.spark.sql.catalyst.expressions.codegen.UnsafeRowWriter(agg_holder, 1); / 039 / agg_result1 = new UnsafeRow(2); / 040 / this.agg_holder1 = new org.apache.spark.sql.catalyst.expressions.codegen.BufferHolder(agg_result1, 0); / 041 / this.agg_rowWriter1 = new org.apache.spark.sql.catalyst.expressions.codegen.UnsafeRowWriter(agg_holder1, 2); / 042 / this.wholestagecodegen_numOutputRows = (org.apache.spark.sql.execution.metric.LongSQLMetric) references[1]; / 043 / wholestagecodegen_metricValue = (org.apache.spark.sql.execution.metric.LongSQLMetricValue) wholestagecodegen_numOutputRows.localValue(); / 044 / } / 045 / / 046 / private void agg_doAggregateWithKeys() throws java.io.IOException { / 047 / /** PRODUCE: INPUT / / 048 / / 049 / while (inputadapter_input.hasNext()) { / 050 / InternalRow inputadapter_row = (InternalRow) inputadapter_input.next(); / 051 / /** CONSUME: TungstenAggregate(key=[id#0L], functions=[(count(1),mode=Final,isDistinct=false)], output=[id#0L,count#4L]) / / 052 / / input[0, bigint] / / 053 / long inputadapter_value = inputadapter_row.getLong(0); / 054 / / input[1, bigint] / / 055 / long inputadapter_value1 = inputadapter_row.getLong(1); / 056 / / 057 / // generate grouping key / 058 / agg_rowWriter.write(0, inputadapter_value); / 059 / / hash(input[0, bigint], 42) / / 060 / int agg_value1 = 42; / 061 / / 062 / agg_value1 = org.apache.spark.unsafe.hash.Murmur3_x86_32.hashLong(inputadapter_value, agg_value1); / 063 / UnsafeRow agg_aggBuffer = null; / 064 / if (true) { / 065 / // try to get the buffer from hash map / 066 / agg_aggBuffer = agg_hashMap.getAggregationBufferFromUnsafeRow(agg_result, agg_value1); / 067 / } / 068 / if (agg_aggBuffer == null) { / 069 / if (agg_sorter == null) { / 070 / agg_sorter = agg_hashMap.destructAndCreateExternalSorter(); / 071 / } else { / 072 / agg_sorter.merge(agg_hashMap.destructAndCreateExternalSorter()); / 073 / } / 074 / / 075 / // the hash map had be spilled, it should have enough memory now, / 076 / // try to allocate buffer again. / 077 / agg_aggBuffer = agg_hashMap.getAggregationBufferFromUnsafeRow(agg_result, agg_value1); / 078 / if (agg_aggBuffer == null) { / 079 / // failed to allocate the first page / 080 / throw new OutOfMemoryError("No enough memory for aggregation"); / 081 / } / 082 / } / 083 / / 084 / // evaluate aggregate function / 085 / / (input[0, bigint] + input[2, bigint]) / / 086 / / input[0, bigint] / / 087 / long agg_value4 = agg_aggBuffer.getLong(0); / 088 / / 089 / long agg_value3 = -1L; / 090 / agg_value3 = agg_value4 + inputadapter_value1; / 091 / // update aggregate buffer / 092 / agg_aggBuffer.setLong(0, agg_value3); / 093 / if (shouldStop()) return; / 094 / } / 095 / / 096 / agg_mapIter = agg_plan.finishAggregate(agg_hashMap, agg_sorter); / 097 / } / 098 / / 099 / protected void processNext() throws java.io.IOException { / 100 / /** PRODUCE: TungstenAggregate(key=[id#0L], functions=[(count(1),mode=Final,isDistinct=false)], output=[id#0L,count#4L]) / / 101 / / 102 / if (!agg_initAgg) { / 103 / agg_initAgg = true; / 104 / agg_doAggregateWithKeys(); / 105 / } / 106 / / 107 / // output the result / 108 / while (agg_mapIter.next()) { / 109 / wholestagecodegen_metricValue.add(1); / 110 / UnsafeRow agg_aggKey = (UnsafeRow) agg_mapIter.getKey(); / 111 / UnsafeRow agg_aggBuffer1 = (UnsafeRow) agg_mapIter.getValue(); / 112 / / 113 / / input[0, bigint] / / 114 / long agg_value6 = agg_aggKey.getLong(0); / 115 / / input[0, bigint] / / 116 / long agg_value7 = agg_aggBuffer1.getLong(0); / 117 / / 118 / /** CONSUME: WholeStageCodegen / / 119 / / 120 / agg_rowWriter1.write(0, agg_value6); / 121 / / 122 / agg_rowWriter1.write(1, agg_value7); / 123 / append(agg_result1); / 124 / / 125 / if (shouldStop()) return; / 126 / } / 127 / / 128 / agg_mapIter.close(); / 129 / if (agg_sorter == null) { / 130 / agg_hashMap.free(); / 131 / } / 132 / } / 133 */ } ``` rxin Author: Eric Liang <ekl@databricks.com> Closes #12025 from ericl/spark-14227.	2016-03-29 13:31:51 -07:00
Dongjoon Hyun	838cb4583d	[MINOR][SQL] Fix exception message to print string-array correctly. ## What changes were proposed in this pull request? This PR is a simple fix for an exception message to print `string[]` content correctly. ```java String[] colPath = requestedSchema.getPaths().get(i); ... - throw new IOException("Required column is missing in data file. Col: " + colPath); + throw new IOException("Required column is missing in data file. Col: " + Arrays.toString(colPath)); ``` ## How was this patch tested? Manual. Author: Dongjoon Hyun <dongjoon@apache.org> Closes #12041 from dongjoon-hyun/fix_exception_message_with_string_array.	2016-03-29 12:47:30 -07:00
Dongjoon Hyun	d612228eff	[MINOR][SQL] Fix typos by replacing 'much' with 'match'. ## What changes were proposed in this pull request? This PR fixes two trivial typos: 'does not much' --> 'does not match'. ## How was this patch tested? Manual. Author: Dongjoon Hyun <dongjoon@apache.org> Closes #12042 from dongjoon-hyun/fix_typo_by_replacing_much_with_match.	2016-03-29 12:45:43 -07:00
Jakob Odersky	d26c42982c	[SPARK-10570][CORE] Add version info to json api Add a new api endpoint `/api/v1/version` to retrieve various version info. This PR only adds support for finding the current spark version, however other version info such as jvm or scala versions can easily be added. Author: Jakob Odersky <jodersky@gmail.com> Closes #10760 from jodersky/SPARK-10570.	2016-03-29 11:10:15 -07:00
Carson Wang	15c0b0006b	[SPARK-14232][WEBUI] Fix event timeline display issue when an executor is removed with a multiple line reason. ## What changes were proposed in this pull request? The event timeline doesn't show on job page if an executor is removed with a multiple line reason. This PR replaces all new line characters in the reason string with spaces. ![timelineerror](https://cloud.githubusercontent.com/assets/9278199/14100211/5fd4cd30-f5be-11e5-9cea-f32651a4cd62.jpg) ## How was this patch tested? Verified on the Web UI. Author: Carson Wang <carson.wang@intel.com> Closes #12029 from carsonwang/eventTimeline.	2016-03-29 11:07:58 -07:00
Yuhao Yang	d2a819a636	[SPARK-14154][MLLIB] Simplify the implementation for Kolmogorov–Smirnov test ## What changes were proposed in this pull request? jira: https://issues.apache.org/jira/browse/SPARK-14154 I just read the code for KolmogorovSmirnovTest and find it could be much simplified following the original definition. Send a PR for discussion ## How was this patch tested? unit test Author: Yuhao Yang <hhbyyh@gmail.com> Closes #11954 from hhbyyh/ksoptimize.	2016-03-29 09:16:50 -07:00
Cheng Lian	a632bb56f8	[SPARK-14208][SQL] Renames spark.sql.parquet.fileScan ## What changes were proposed in this pull request? Renames SQL option `spark.sql.parquet.fileScan` since now all `HadoopFsRelation` based data sources are being migrated to `FileScanRDD` code path. ## How was this patch tested? None. Author: Cheng Lian <lian@databricks.com> Closes #12003 from liancheng/spark-14208-option-renaming.	2016-03-29 20:56:01 +08:00
Bryan Cutler	425bcf6d68	[SPARK-13963][ML] Adding binary toggle param to HashingTF ## What changes were proposed in this pull request? Adding binary toggle parameter to ml.feature.HashingTF, as well as mllib.feature.HashingTF since the former wraps this functionality. This parameter, if true, will set non-zero valued term counts to 1 to transform term count features to binary values that are well suited for discrete probability models. ## How was this patch tested? Added unit tests for ML and MLlib Author: Bryan Cutler <cutlerb@gmail.com> Closes #11832 from BryanCutler/binary-param-HashingTF-SPARK-13963.	2016-03-29 12:30:30 +02:00
Wenchen Fan	83775bc78e	[SPARK-14158][SQL] implement buildReader for json data source ## What changes were proposed in this pull request? This PR implements buildReader for json data source and enable it in the new data source code path. ## How was this patch tested? existing tests Author: Wenchen Fan <wenchen@databricks.com> Closes #11960 from cloud-fan/json.	2016-03-29 14:34:12 +08:00
wm624@hotmail.com	63b200e8d4	[SPARK-14071][PYSPARK][ML] Change MLWritable.write to be a property Add property to MLWritable.write method, so we can use .write instead of .write() Add a new test to ml/test.py to check whether the write is a property. ./python/run-tests --python-executables=python2.7 --modules=pyspark-ml Will test against the following Python executables: ['python2.7'] Will test the following Python modules: ['pyspark-ml'] Finished test(python2.7): pyspark.ml.evaluation (11s) Finished test(python2.7): pyspark.ml.clustering (16s) Finished test(python2.7): pyspark.ml.classification (24s) Finished test(python2.7): pyspark.ml.recommendation (24s) Finished test(python2.7): pyspark.ml.feature (39s) Finished test(python2.7): pyspark.ml.regression (26s) Finished test(python2.7): pyspark.ml.tuning (15s) Finished test(python2.7): pyspark.ml.tests (30s) Tests passed in 55 seconds Author: wm624@hotmail.com <wm624@hotmail.com> Closes #11945 from wangmiao1981/fix_property.	2016-03-28 22:33:25 -07:00
sethah	f6066b0c3c	[SPARK-11730][ML] Add feature importances for GBTs. ## What changes were proposed in this pull request? Now that GBTs have been moved to ML, they can use the implementation of feature importance for random forests. This patch simply adds a `featureImportances` attribute to `GBTClassifier` and `GBTRegressor` and adds tests for each. GBT feature importances here simply average the feature importances for each tree in its ensemble. This follows the implementation from scikit-learn. This method is also suggested by J Friedman in [this paper](https://statweb.stanford.edu/~jhf/ftp/trebst.pdf). ## How was this patch tested? Unit tests were added to `GBTClassifierSuite` and `GBTRegressorSuite` to validate feature importances. Author: sethah <seth.hendrickson16@gmail.com> Closes #11961 from sethah/SPARK-11730.	2016-03-28 22:27:53 -07:00
Sun Rui	d3638d7bff	[SPARK-12792] [SPARKR] Refactor RRDD to support R UDF. ## What changes were proposed in this pull request? Refactor RRDD by separating the common logic interacting with the R worker to a new class RRunner, which can be used to evaluate R UDFs. Now RRDD relies on RRuner for RDD computation and RRDD could be reomved if we want to remove RDD API in SparkR later. ## How was this patch tested? dev/lint-r SparkR unit tests Author: Sun Rui <rui.sun@intel.com> Closes #12024 from sun-rui/SPARK-12792_new.	2016-03-28 21:51:02 -07:00

... 2 3 4 5 6 ...

15547 commits