ODIn/spark-instrumented-optimizer

Author	SHA1	Message	Date
Mark Hamstra	51458ab4a1	Added stageId <--> jobId mapping in DAGScheduler ...and make sure that DAGScheduler data structures are cleaned up on job completion. Initial effort and discussion at https://github.com/mesos/spark/pull/842	2013-12-03 09:57:31 -08:00
Harvey Feng	46b87b8a25	Merge pull request #2 from colorant/yarn-client-2.2 Fix pom.xml for maven build	2013-12-03 00:41:11 -08:00
Raymond Liu	4738818dd6	Fix pom.xml for maven build	2013-12-03 16:36:05 +08:00
Harvey Feng	1b6e450771	Use published "org.spark-project.akka-*" in sbt build for Hadoop-2.2 dependencies. This also includes: -Change `isNewYarn` to `isNewHadoop`, since the protobuf-2.5 dependency is from Hadoop-2.2 itself. -Regexp bugix Credits to @alig for this patch.	2013-12-03 00:28:33 -08:00
Reynold Xin	58d9bbcfec	Merge pull request #217 from aarondav/mesos-urls Re-enable zk:// urls for Mesos SparkContexts This was broken in PR #71 when we explicitly disallow anything that didn't fit a mesos:// url. Although it is not really clear that a zk:// url should match Mesos, it is what the docs say and it is necessary for backwards compatibility. Additionally added a unit test for the creation of all types of TaskSchedulers. Since YARN and Mesos are not necessarily available in the system, they are allowed to pass as long as the YARN/Mesos code paths are exercised.	2013-12-02 21:58:53 -08:00
Prashant Sharma	09e8be9a62	Made running SparkActorSystem specific to executors only.	2013-12-03 11:27:45 +05:30
Aaron Davidson	0f24576c08	Cleanup and documentation of SparkActorSystem	2013-12-03 11:05:12 +05:30
Reynold Xin	e34b4693d3	Mark partitioner, name, and generator field in RDD as @transient.	2013-12-02 21:24:44 -08:00
Kay Ousterhout	58b3aff9a8	Fixed problem with scheduler delay	2013-12-02 20:30:03 -08:00
Aaron Davidson	f6c8c1c7b6	Cleanup and documentation of SparkActorSystem	2013-12-02 11:42:53 -08:00
Prashant Sharma	5b11028a04	Made akka capable of tolerating fatal exceptions and moving on.	2013-12-02 10:47:39 +05:30
Reynold Xin	740922f25d	Merge pull request #219 from sundeepn/schedulerexception Scheduler quits when newStage fails The current scheduler thread does not handle exceptions from newStage stage while launching new jobs. The thread fails on any exception that gets triggered at that level, leaving the cluster hanging with no schduler.	2013-12-01 12:46:58 -08:00
Sundeep Narravula	be3ea2394f	Log exception in scheduler in addition to passing it to the caller. Code Styling changes.	2013-12-01 00:50:34 -08:00
Reynold Xin	60e23a58b2	Merge pull request #216 from liancheng/fix-spark-966 Bugfix: SPARK-965 & SPARK-966 SPARK-965: https://spark-project.atlassian.net/browse/SPARK-965 SPARK-966: https://spark-project.atlassian.net/browse/SPARK-966 * Add back `DAGScheduler.start()`, `eventProcessActor` is created and started here. Notice that function is only called by `SparkContext`. * Cancel the scheduled stage resubmission task when stopping `eventProcessActor` * Add a new `DAGSchedulerEvent` `ResubmitFailedStages` This event message is sent by the scheduled stage resubmission task to `eventProcessActor`. In this way, `DAGScheduler.resubmitFailedStages()` is guaranteed to be executed from the same thread that runs `DAGScheduler.processEvent()`. Please refer to discussion in [SPARK-966](https://spark-project.atlassian.net/browse/SPARK-966) for details.	2013-11-30 23:38:49 -08:00
Reynold Xin	9cf7f31e4d	Memoize preferred locations in ZippedPartitionsBaseRDD so preferred location computation doesn't lead to exponential explosion. (cherry picked from commit `e36fe55a03`) Signed-off-by: Reynold Xin <rxin@apache.org>	2013-11-30 18:10:52 -08:00
Sundeep Narravula	4d53830eb7	Scheduler quits when createStage fails. The current scheduler thread does not handle exceptions from createStage stage while launching new jobs. The thread fails on any exception that gets triggered at that level, leaving the cluster hanging with no schduler.	2013-11-30 16:18:12 -08:00
Aaron Davidson	96df26be47	Add spaces between tests	2013-11-29 13:20:43 -08:00
Prashant Sharma	5618af6803	Merge branch 'master' into wip-scala-2.10	2013-11-29 13:41:21 +05:30
Prashant Sharma	1bc83ca791	Changed defaults for akka to almost disable failure detector.	2013-11-29 13:41:05 +05:30
Lian, Cheng	4a1d966e26	More comments	2013-11-29 16:02:58 +08:00
Lian, Cheng	1e25086009	Updated some inline comments in DAGScheduler	2013-11-29 15:56:47 +08:00
Josh Rosen	3787f514d9	Fix UnicodeEncodeError in PySpark saveAsTextFile(). Fixes SPARK-970.	2013-11-28 23:44:56 -08:00
Aaron Davidson	081a0b6861	Add unit test for SparkContext scheduler creation Since YARN and Mesos are not necessarily available in the system, they are allowed to pass as long as the YARN/Mesos code paths are exercised.	2013-11-28 20:40:57 -08:00
Aaron Davidson	37f161cf6b	Re-enable zk:// urls for Mesos SparkContexts This was broken in PR #71 when we explicitly disallow anything that didn't fit a mesos:// url. Although it is not really clear that a zk:// url should match Mesos, it is what the docs say and it is necessary for backwards compatibility.	2013-11-28 20:37:56 -08:00
Lian, Cheng	18def5d6f2	Bugfix: SPARK-965 & SPARK-966 SPARK-965: https://spark-project.atlassian.net/browse/SPARK-965 SPARK-966: https://spark-project.atlassian.net/browse/SPARK-966 * Add back DAGScheduler.start(), eventProcessActor is created and started here. Notice that function is only called by SparkContext. * Cancel the scheduled stage resubmission task when stopping eventProcessActor * Add a new DAGSchedulerEvent ResubmitFailedStages This event message is sent by the scheduled stage resubmission task to eventProcessActor. In this way, DAGScheduler.resubmitFailedStages is guaranteed to be executed from the same thread that runs DAGScheduler.processEvent. Please refer to discussion in SPARK-966 for details.	2013-11-28 17:46:06 +08:00
Prashant Sharma	3ec5d74766	Fixed the broken build.	2013-11-28 13:02:28 +05:30
Matei Zaharia	743a31a7ca	Merge pull request #210 from haitaoyao/http-timeout add http timeout for httpbroadcast While pulling task bytecode from HttpBroadcast server, there's no timeout value set. This may cause spark executor code hang and other task in the same executor process wait for the lock. I have encountered the issue in my cluster. Here's the stacktrace I captured : https://gist.github.com/haitaoyao/7655830 So add a time out value to ensure the task fail fast.	2013-11-27 18:24:39 -08:00
Prashant Sharma	17987778da	Merge branch 'master' into wip-scala-2.10 Conflicts: core/src/main/scala/org/apache/spark/api/python/PythonRDD.scala core/src/main/scala/org/apache/spark/rdd/MapPartitionsRDD.scala core/src/main/scala/org/apache/spark/rdd/MapPartitionsWithContextRDD.scala core/src/main/scala/org/apache/spark/rdd/RDD.scala python/pyspark/rdd.py	2013-11-27 14:44:12 +05:30
Harvey Feng	993e293d6e	Merge pull request #1 from colorant/yarn-client-2.2 Port yarn-client mode for new-yarn	2013-11-27 00:57:54 -08:00
Prashant Sharma	54862af5ee	Improvements from the review comments and followed Boy Scout Rule.	2013-11-27 14:26:28 +05:30
Raymond Liu	403cac9be3	Port yarn-client mode for new-yarn	2013-11-27 16:10:42 +08:00
Matei Zaharia	fb6875dd5c	Merge pull request #146 from JoshRosen/pyspark-custom-serializers Custom Serializers for PySpark This pull request adds support for custom serializers to PySpark. For now, all Python-transformed (or parallelize()d RDDs) are serialized with the same serializer that's specified when creating SparkContext. For now, PySpark includes `PickleSerDe` and `MarshalSerDe` classes for using Python's `pickle` and `marshal` serializers. It's pretty easy to add support for other serializers, although I still need to add instructions on this. A few notable changes: - The Scala `PythonRDD` class no longer manipulates Pickled objects; data from `textFile` is written to Python as MUTF-8 strings. The Python code performs the appropriate bookkeeping to track which deserializer should be used when reading an underlying JavaRDD. This mechanism could also be used to support other data exchange formats, such as MsgPack. - Several magic numbers were refactored into constants. - Batching is implemented by wrapping / decorating an unbatched SerDe.	2013-11-26 20:55:40 -08:00
Matei Zaharia	330ada1766	Merge pull request #207 from henrydavidge/master Log a warning if a task's serialized size is very big As per Reynold's instructions, we now create a warning level log entry if a task's serialized size is too big. "Too big" is currently defined as 100kb. This warning message is generated at most once for each stage.	2013-11-26 19:08:33 -08:00
Matei Zaharia	615213fb82	Merge pull request #212 from markhamstra/SPARK-963 [SPARK-963] Fixed races in JobLoggerSuite	2013-11-26 19:07:20 -08:00
Harvey Feng	afe4fe7f5e	Merge remote-tracking branch 'origin/master' into yarn-2.2 Conflicts: yarn/src/main/scala/org/apache/spark/deploy/yarn/Client.scala	2013-11-26 15:03:03 -08:00
Harvey Feng	a1a1c62a3e	Add optional Hadoop 2.2 settings in sbt build. If the Hadoop used is version 2.2 or derived from it, then Spark will be compiled against protobuf-2.5 and a protobuf-2.5 version of Akka 2.0.5.	2013-11-26 14:58:41 -08:00
Josh Rosen	1b74a27da0	Removed unused basestring case from dump_stream.	2013-11-26 14:35:12 -08:00
hhd	57579934f0	Emit warning when task size > 100KB	2013-11-26 16:58:39 -05:00
Mark Hamstra	ed7ecb93ce	[SPARK-963] Wait for SparkListenerBus eventQueue to be empty before checking jobLogger state	2013-11-26 13:30:17 -08:00
Reynold Xin	cb976dfb50	Merge pull request #209 from pwendell/better-docs Improve docs for shuffle instrumentation	2013-11-26 10:23:19 -08:00
Prashant Sharma	dca946ff67	Documenting the newly added spark properties.	2013-11-26 20:47:38 +05:30
Prashant Sharma	560e44a8e1	Restored master address for client.	2013-11-26 18:18:05 +05:30
haitao.yao	db998a6e14	add http timeout for httpbroadcast	2013-11-26 18:23:48 +08:00
Prashant Sharma	d092a8cc6a	Fixed compile time warnings and formatting post merge.	2013-11-26 15:21:50 +05:30
Matei Zaharia	18d6df0e17	Merge pull request #86 from holdenk/master Add histogram functionality to DoubleRDDFunctions This pull request add histogram functionality to the DoubleRDDFunctions.	2013-11-26 00:00:07 -08:00
Patrick Wendell	297c09d4bb	Improve docs for shuffle instrumentation	2013-11-25 22:53:28 -08:00
Holden Karau	7222ee2977	Fix the test	2013-11-25 21:06:42 -08:00
Matei Zaharia	0e2109ddb2	Merge pull request #204 from rxin/hash OpenHashSet fixes Incorporated ideas from pull request #200. - Use Murmur Hash 3 finalization step to scramble the bits of HashCode instead of the simpler version in java.util.HashMap; the latter one had trouble with ranges of consecutive integers. Murmur Hash 3 is used by fastutil. - Don't check keys for equality when re-inserting due to growing the table; the keys will already be unique. - Remember the grow threshold instead of recomputing it on each insert Also added unit tests for size estimation for specialized hash sets and maps.	2013-11-25 20:48:37 -08:00
Matei Zaharia	c46067f096	Merge pull request #206 from ash211/patch-2 Update tuning.md Clarify when serializer is used based on recent user@ mailing list discussion.	2013-11-25 19:09:31 -08:00
Matei Zaharia	14bb465bb3	Merge pull request #201 from rxin/mappartitions Use the proper partition index in mapPartitionsWIthIndex mapPartitionsWithIndex uses TaskContext.partitionId as the partition index. TaskContext.partitionId used to be identical to the partition index in a RDD. However, pull request #186 introduced a scenario (with partition pruning) that the two can be different. This pull request uses the right partition index in all mapPartitionsWithIndex related calls. Also removed the extra MapPartitionsWIthContextRDD and put all the mapPartitions related functionality in MapPartitionsRDD.	2013-11-25 18:50:18 -08:00

... 10 11 12 13 14 ...

5335 commits