ODIn/spark-instrumented-optimizer

Author	SHA1	Message	Date
Harvey Feng	993e293d6e	Merge pull request #1 from colorant/yarn-client-2.2 Port yarn-client mode for new-yarn	2013-11-27 00:57:54 -08:00
Raymond Liu	403cac9be3	Port yarn-client mode for new-yarn	2013-11-27 16:10:42 +08:00
Harvey Feng	afe4fe7f5e	Merge remote-tracking branch 'origin/master' into yarn-2.2 Conflicts: yarn/src/main/scala/org/apache/spark/deploy/yarn/Client.scala	2013-11-26 15:03:03 -08:00
Harvey Feng	a1a1c62a3e	Add optional Hadoop 2.2 settings in sbt build. If the Hadoop used is version 2.2 or derived from it, then Spark will be compiled against protobuf-2.5 and a protobuf-2.5 version of Akka 2.0.5.	2013-11-26 14:58:41 -08:00
Reynold Xin	cb976dfb50	Merge pull request #209 from pwendell/better-docs Improve docs for shuffle instrumentation	2013-11-26 10:23:19 -08:00
Matei Zaharia	18d6df0e17	Merge pull request #86 from holdenk/master Add histogram functionality to DoubleRDDFunctions This pull request add histogram functionality to the DoubleRDDFunctions.	2013-11-26 00:00:07 -08:00
Patrick Wendell	297c09d4bb	Improve docs for shuffle instrumentation	2013-11-25 22:53:28 -08:00
Holden Karau	7222ee2977	Fix the test	2013-11-25 21:06:42 -08:00
Matei Zaharia	0e2109ddb2	Merge pull request #204 from rxin/hash OpenHashSet fixes Incorporated ideas from pull request #200. - Use Murmur Hash 3 finalization step to scramble the bits of HashCode instead of the simpler version in java.util.HashMap; the latter one had trouble with ranges of consecutive integers. Murmur Hash 3 is used by fastutil. - Don't check keys for equality when re-inserting due to growing the table; the keys will already be unique. - Remember the grow threshold instead of recomputing it on each insert Also added unit tests for size estimation for specialized hash sets and maps.	2013-11-25 20:48:37 -08:00
Matei Zaharia	c46067f096	Merge pull request #206 from ash211/patch-2 Update tuning.md Clarify when serializer is used based on recent user@ mailing list discussion.	2013-11-25 19:09:31 -08:00
Matei Zaharia	14bb465bb3	Merge pull request #201 from rxin/mappartitions Use the proper partition index in mapPartitionsWIthIndex mapPartitionsWithIndex uses TaskContext.partitionId as the partition index. TaskContext.partitionId used to be identical to the partition index in a RDD. However, pull request #186 introduced a scenario (with partition pruning) that the two can be different. This pull request uses the right partition index in all mapPartitionsWithIndex related calls. Also removed the extra MapPartitionsWIthContextRDD and put all the mapPartitions related functionality in MapPartitionsRDD.	2013-11-25 18:50:18 -08:00
Andrew Ash	08afef37a0	Update tuning.md Clarify when serializer is used based on recent user@ mailing list discussion.	2013-11-25 17:08:52 -08:00
Matei Zaharia	eb4296c8f7	Merge pull request #101 from colorant/yarn-client-scheduler For SPARK-527, Support spark-shell when running on YARN sync to trunk and resubmit here In current YARN mode approaching, the application is run in the Application Master as a user program thus the whole spark context is on remote. This approaching won't support application that involve local interaction and need to be run on where it is launched. So In this pull request I have a YarnClientClusterScheduler and backend added. With this scheduler, the user application is launched locally,While the executor will be launched by YARN on remote nodes with a thin AM which only launch the executor and monitor the Driver Actor status, so that when client app is done, it can finish the YARN Application as well. This enables spark-shell to run upon YARN. This also enable other Spark applications to have the spark context to run locally with a master-url "yarn-client". Thus e.g. SparkPi could have the result output locally on console instead of output in the log of the remote machine where AM is running on. Docs also updated to show how to use this yarn-client mode.	2013-11-25 15:25:29 -08:00
Reynold Xin	466fd06475	Incorporated ideas from pull request #200 . - Use Murmur Hash 3 finalization step to scramble the bits of HashCode instead of the simpler version in java.util.HashMap; the latter one had trouble with ranges of consecutive integers. Murmur Hash 3 is used by fastutil. - Don't check keys for equality when re-inserting due to growing the table; the keys will already be unique - Remember the grow threshold instead of recomputing it on each insert	2013-11-25 18:27:26 +08:00
Reynold Xin	95c55df1c2	Added unit tests for size estimation for specialized hash sets and maps.	2013-11-25 18:27:06 +08:00
Reynold Xin	62889c419c	Merge pull request #203 from witgo/master Fix Maven build for metrics-graphite	2013-11-25 11:27:45 +08:00
LiGuoqiang	989203604e	Fix Maven build for metrics-graphite	2013-11-25 11:23:11 +08:00
Matei Zaharia	859d62dc2a	Merge pull request #151 from russellcardullo/add-graphite-sink Add graphite sink for metrics This adds a metrics sink for graphite. The sink must be configured with the host and port of a graphite node and optionally may be configured with a prefix that will be prepended to all metrics that are sent to graphite.	2013-11-24 16:19:51 -08:00
Matei Zaharia	65de73c7f8	Merge pull request #185 from mkolod/random-number-generator XORShift RNG with unit tests and benchmark This patch was introduced to address SPARK-950 - the discussion below the ticket explains not only the rationale, but also the design and testing decisions: https://spark-project.atlassian.net/browse/SPARK-950 To run unit test, start SBT console and type: compile test-only org.apache.spark.util.XORShiftRandomSuite To run benchmark, type: project core console Once the Scala console starts, type: org.apache.spark.util.XORShiftRandom.benchmark(100000000) XORShiftRandom is also an object with a main method taking the number of iterations as an argument, so you can also run it from the command line.	2013-11-24 15:52:33 -08:00
Reynold Xin	972171b9d9	Merge pull request #197 from aarondav/patrick-fix Fix 'timeWriting' stat for shuffle files Due to concurrent git branches, changes from shuffle file consolidation patch caused the shuffle write timing patch to no longer actually measure the time, since it requires time be measured after the stream has been closed.	2013-11-25 07:50:46 +08:00
Reynold Xin	e9ff13ec72	Consolidated both mapPartitions related RDDs into a single MapPartitionsRDD. Also changed the semantics of the index parameter in mapPartitionsWithIndex from the partition index of the output partition to the partition index in the current RDD.	2013-11-24 17:56:43 +08:00
Reynold Xin	718cc803f7	Merge pull request #200 from mateiz/hash-fix AppendOnlyMap fixes - Chose a more random reshuffling step for values returned by Object.hashCode to avoid some long chaining that was happening for consecutive integers (e.g. `sc.makeRDD(1 to 100000000, 100).map(t => (t, t)).reduceByKey(_ + _).count`) - Some other small optimizations throughout (see commit comments)	2013-11-24 11:02:02 +08:00
Matei Zaharia	9837a60234	Some other optimizations to AppendOnlyMap: - Don't check keys for equality when re-inserting due to growing the table; the keys will already be unique - Remember the grow threshold instead of recomputing it on each insert	2013-11-23 17:38:29 -08:00
Matei Zaharia	7535d7fbcb	Fixes to AppendOnlyMap: - Use Murmur Hash 3 finalization step to scramble the bits of HashCode instead of the simpler version in java.util.HashMap; the latter one had trouble with ranges of consecutive integers. Murmur Hash 3 is used by fastutil. - Use Object.equals() instead of Scala's == to compare keys, because the latter does extra casts for numeric types (see the equals method in https://github.com/scala/scala/blob/master/src/library/scala/runtime/BoxesRunTime.java)	2013-11-23 17:21:37 -08:00
Harvey Feng	4f1c3fa5d7	Hadoop 2.2 YARN API migration for `SPARK_HOME/new-yarn`	2013-11-23 17:08:30 -08:00
Harvey Feng	ab8652f2d3	Add a "new-yarn" directory in SPARK_HOME, intended to contain Hadoop-2.2 API changes.	2013-11-23 17:08:30 -08:00
Harvey Feng	a67ebf4377	A few more style fixes in `yarn` package.	2013-11-23 17:08:30 -08:00
Reynold Xin	51aa9d6e99	Merge pull request #198 from ankurdave/zipPartitions-preservesPartitioning Support preservesPartitioning in RDD.zipPartitions In `RDD.zipPartitions`, add support for a `preservesPartitioning` option (similar to `RDD.mapPartitions`) that reuses the first RDD's partitioner.	2013-11-23 19:46:46 +08:00
Ankur Dave	c1507afc6c	Support preservesPartitioning in RDD.zipPartitions	2013-11-23 03:03:31 -08:00
Aaron Davidson	ccea38b759	Fix 'timeWriting' stat for shuffle files Due to concurrent git branches, changes from shuffle file consolidation patch caused the shuffle write timing patch to no longer actually measure the time, since it requires time be measured after the stream has been closed.	2013-11-21 21:36:08 -08:00
Reynold Xin	086b097e33	Merge pull request #193 from aoiwelle/patch-1 Fix Kryo Serializer buffer documentation inconsistency The documentation here is inconsistent with the coded default and other documentation.	2013-11-22 10:26:39 +08:00
Reynold Xin	f20093c3af	Merge pull request #196 from pwendell/master TimeTrackingOutputStream should pass on calls to close() and flush(). Without this fix you get a huge number of open files when running shuffles.	2013-11-22 10:12:13 +08:00
Raymond Liu	ab3cefde53	Add YarnClientClusterScheduler and Backend. With this scheduler, the user application is launched locally, While the executor will be launched by YARN on remote nodes. This enables spark-shell to run upon YARN.	2013-11-22 09:23:27 +08:00
Patrick Wendell	53b94ef2f5	TimeTrackingOutputStream should pass on calls to close() and flush(). Without this fix you get a huge number of open shuffles after running shuffles.	2013-11-21 17:20:15 -08:00
Harvey Feng	9eae80f111	Merge branch 'master' into yarn-cleanup Conflicts: yarn/src/main/scala/org/apache/spark/deploy/yarn/ApplicationMaster.scala yarn/src/main/scala/org/apache/spark/deploy/yarn/Client.scala yarn/src/main/scala/org/apache/spark/deploy/yarn/WorkerRunnable.scala yarn/src/main/scala/org/apache/spark/deploy/yarn/YarnAllocationHandler.scala	2013-11-21 03:41:57 -08:00
Neal Wiggins	21b5478ed6	Fix Kryo Serializer buffer inconsistency The documentation here is inconsistent with the coded default and other documentation.	2013-11-20 16:19:25 -08:00
Reynold Xin	2fead510f7	Merge branch 'master' of github.com:tbfenet/incubator-spark PartitionPruningRDD is using index from parent I was getting a ArrayIndexOutOfBoundsException exception after doing union on pruned RDD. The index it was using on the partition was the index in the original RDD not the new pruned RDD.	2013-11-21 07:15:55 +08:00
Matei Zaharia	4b895013cc	Merge pull request #191 from hsaputra/removesemicolonscala Cleanup to remove semicolons (;) from Scala code -) The main reason for this PR is to remove semicolons from single statements of Scala code. -) Remove unused imports as I see them -) Fix ASF comment header from some of files (bad copy paste I suppose)	2013-11-20 10:36:10 -08:00
Marek Kolodziej	22724659db	Make XORShiftRandom explicit in KMeans and roll it back for RDD	2013-11-20 07:03:36 -05:00
Marek Kolodziej	bcc6ed30bf	Formatting and scoping (private[spark]) updates	2013-11-19 20:50:38 -05:00
Henry Saputra	43dfac5132	Merge branch 'master' into removesemicolonscala	2013-11-19 16:57:57 -08:00
Henry Saputra	10be58f251	Another set of changes to remove unnecessary semicolon (;) from Scala code. Passed the sbt/sbt compile and test	2013-11-19 16:56:23 -08:00
Matei Zaharia	f568912f85	Merge pull request #181 from BlackNiuza/fix_tasks_number correct number of tasks in ExecutorsUI Index `a` is not `execId` here	2013-11-19 16:11:31 -08:00
Matei Zaharia	aa638ed9c1	Merge pull request #189 from tgravescs/sparkYarnErrorHandling Impove Spark on Yarn Error handling Improve cli error handling and only allow a certain number of worker failures before failing the application. This will help prevent users from doing foolish things and their jobs running forever. For instance using 32 bit java but trying to allocate 8G containers. This loops forever without this change, now it errors out after a certain number of retries. The number of tries is configurable. Also increase the frequency we ping the RM to increase speed at which we get containers if they die. The Yarn MR app defaults to pinging the RM every 1 seconds, so the default of 5 seconds here is fine. But that is configurable as well in case people want to change it. I do want to make sure there aren't any cases that calling stopExecutors in CoarseGrainedSchedulerBackend would cause problems? I couldn't think of any and testing on standalone cluster as well as yarn.	2013-11-19 16:05:44 -08:00
Matei Zaharia	55925805fc	Merge pull request #187 from aarondav/example-bcast-test Enable the Broadcast examples to work in a cluster setting Since they rely on println to display results, we need to first collect those results to the driver to have them actually display locally. This issue came up on the mailing lists [here](http://mail-archives.apache.org/mod_mbox/incubator-spark-user/201311.mbox/%3C2013111909591557147628%40ict.ac.cn%3E).	2013-11-19 16:04:01 -08:00
tgravescs	4093e9393a	Impove Spark on Yarn Error handling	2013-11-19 12:44:00 -06:00
Henry Saputra	9c934b640f	Remove the semicolons at the end of Scala code to make it more pure Scala code. Also remove unused imports as I found them along the way. Remove return statements when returning value in the Scala code. Passing compile and tests.	2013-11-19 10:19:03 -08:00
Matthew Taylor	f639b65eab	PartitionPruningRDD is using index from parent(review changes)	2013-11-19 10:48:48 +00:00
Aaron Davidson	50fd8d98c0	Enable the Broadcast examples to work in a cluster setting Since they rely on println to display results, we need to first collect those results to the driver to have them actually display locally.	2013-11-18 22:51:35 -08:00
Matthew Taylor	13b9bf494b	PartitionPruningRDD is using index from parent	2013-11-19 06:27:33 +00:00

1 2 3 4 5 ...

4613 commits