ODIn/spark-instrumented-optimizer

Author	SHA1	Message	Date
Grace Huang	892fb8ffa8	remedy the line-wrap while exceeding 100 chars	2013-09-30 20:12:55 +08:00
Harvey Feng	7d06bdde1d	Merge HadoopDatasetRDD into HadoopRDD.	2013-09-29 20:08:03 -07:00
Grace Huang	4b68be5f3c	SPARK-900 Use coarser grained naming for metrics	2013-09-27 14:47:38 +08:00
Harvey Feng	417085716a	Merge remote-tracking branch 'oldsparkme/hadoopRDD-broadcast-change' into hadoop-config-cache	2013-09-26 15:49:42 -07:00
Aaron Davidson	42d72308fb	Add license notices	2013-09-26 15:45:20 -07:00
Aaron Davidson	f549ea33d3	Standalone Scheduler fault tolerance using ZooKeeper This patch implements full distributed fault tolerance for standalone scheduler Masters. There is only one master Leader at a time, which is actively serving scheduling requests. If this Leader crashes, another master will eventually be elected, reconstruct the state from the first Master, and continue serving scheduling requests. Leader election is performed using the ZooKeeper leader election pattern. We try to minimize the use of ZooKeeper and the assumptions about ZooKeeper's behavior, so there is a layer of retries and session monitoring on top of the ZooKeeper client. Master failover follows directly from the single-node Master recovery via the file system (patch 194ba4b8), save that the Master state is stored in ZooKeeper instead. Configuration: By default, no recovery mechanism is enabled (spark.deploy.recoveryMode = NONE). By setting spark.deploy.recoveryMode to ZOOKEEPER and setting spark.deploy.zookeeper.url to an appropriate ZooKeeper URL, ZooKeeper recovery mode is enabled. By setting spark.deploy.recoveryMode to FILESYSTEM and setting spark.deploy.recoveryDirectory to an appropriate directory accessible by the Master, we will keep the behavior of from 194ba4b8. Additionally, places where a Master could be specificied by a spark:// url can now take comma-delimited lists to specify backup masters. Note that this is only used for registration of NEW Workers and application Clients. Once a Worker or Client has registered with the Master Leader, it is "in the system" and will never need to register again. Forthcoming: Documentation, tests (! - only ad hoc testing has been performed so far) I do not intend for this commit to be merged until tests are added, but this patch should still be mostly reviewable until then.	2013-09-26 15:04:23 -07:00
Aaron Davidson	d5a96feccb	Standalone Scheduler fault recovery Implements a basic form of Standalone Scheduler fault recovery. In particular, this allows faults to be manually recovered from by means of restarting the Master process on the same machine. This is the majority of the code necessary for general fault tolerance, which will first elect a leader and then recover the Master state. In order to enable fault recovery, the Master will persist a small amount of state related to the registration of Workers and Applications to disk. If the Master is started and sees that this state is still around, it will enter Recovery mode, during which time it will not schedule any new Executors on Workers (but it does accept the registration of new Clients and Workers). At this point, the Master attempts to reconnect to all Workers and Client applications that were registered at the time of failure. After confirming either the existence or nonexistence of all such nodes (within a certain timeout), the Master will exit Recovery mode and resume normal scheduling.	2013-09-26 14:59:35 -07:00
Reynold Xin	714fdabd99	Merge pull request #17 from rxin/optimize Remove -optimize flag	2013-09-26 14:28:55 -07:00
Reynold Xin	13eced723f	Merge pull request #16 from pwendell/master Bug fix in master build	2013-09-26 14:18:19 -07:00
Reynold Xin	70a0b993d4	Merge pull request #14 from kayousterhout/untangle_scheduler Improved organization of scheduling packages. This commit does not change any code -- only file organization. Please let me know if there was some masterminded strategy behind the existing organization that I failed to understand! There are two components of this change: (1) Moving files out of the cluster package, and down a level to the scheduling package. These files are all used by the local scheduler in addition to the cluster scheduler(s), so should not be in the cluster package. As a result of this change, none of the files in the local package reference files in the cluster package. (2) Moving the mesos package to within the cluster package. The mesos scheduling code is for a cluster, and represents a specific case of cluster scheduling (the Mesos-related classes often subclass cluster scheduling classes). Thus, the most logical place for it seems to be within the cluster package. The one thing about the scheduling code that seems a little funny to me is the naming of the SchedulerBackends. The StandaloneSchedulerBackend is not just for Standalone mode, but instead is used by Mesos coarse grained mode and Yarn, and the backend that is just for Standalone mode is instead called SparkDeploySchedulerBackend. I didn't change this because I wasn't sure if there was a reason for this naming that I'm just not aware of.	2013-09-26 14:11:54 -07:00
Reynold Xin	76677b8fa1	Merge pull request #670 from jey/ec2-ssh-improvements EC2 SSH improvements	2013-09-26 14:03:46 -07:00
Reynold Xin	3f283278b0	Removed scala -optimize flag.	2013-09-26 13:58:10 -07:00
Reynold Xin	c514cd1587	Merge pull request #930 from holdenk/master Add mapPartitionsWithIndex	2013-09-26 13:48:20 -07:00
Patrick Wendell	e2ff59af72	Bug fix in master build	2013-09-26 13:06:51 -07:00
Reynold Xin	560ee5c9bb	Merge pull request #7 from wannabeast/memorystore-fixes some minor fixes to MemoryStore This is a repeat of #5, moved to its own branch in my repo. This makes all updates to on ; it skips on synchronizing the reads where it can get away with it.	2013-09-26 11:27:34 -07:00
Patrick Wendell	6566a19b38	Merge pull request #9 from rxin/limit Smarter take/limit implementation.	2013-09-26 08:01:04 -07:00
Kay Ousterhout	d85fe41b2b	Improved organization of scheduling packages. This commit does not change any code -- only file organization. There are two components of this change: (1) Moving files out of the cluster package, and down a level to the scheduling package. These files are all used by the local scheduler in addition to the cluster scheduler(s), so should not be in the cluster package. As a result of this change, none of the files in the local package reference files in the cluster package. (2) Moving the mesos package to within the cluster package. The mesos scheduling code is for a cluster, and represents a specific case of cluster scheduling (the Mesos-related classes often subclass cluster scheduling classes). Thus, the most logical place for it is within the cluster package.	2013-09-25 12:45:46 -07:00
Patrick Wendell	9d34838bde	Merge remote-tracking branch 'apache-github/pr/13' into HEAD	2013-09-24 15:31:12 -07:00
Patrick Wendell	6079721fa1	Update build version in master	2013-09-24 11:41:51 -07:00
Holden Karau	0cef683553	Fix formatting :)	2013-09-23 19:39:42 -07:00
Reynold Xin	7220e8f90b	Merge remote-tracking branch 'pr/12' Fix spacing so java.io.tmpdir doesn't run on with SPARK_JAVA_OPTS	2013-09-23 14:02:21 -07:00
$Y.CORP.YAHOO.COM\tgraves$ Y.CORP.YAHOO.COM\tgraves	a314b30733	Fix spacing so that the java.io.tmpdir doesn't run on with SPARK_JAVA_OPTS	2013-09-23 14:48:17 -05:00
Reynold Xin	0d2e5c3e24	Merge branch 'master' of https://git-wip-us.apache.org/repos/asf/incubator-spark	2013-09-23 11:55:55 -07:00
Reynold Xin	ff540a015b	Merge branch 'master' of github.com:markhamstra/incubator-spark	2013-09-23 11:55:02 -07:00
Reynold Xin	f4dc9d37f8	Merge branch 'master' of github.com:mesos/spark	2013-09-23 11:52:52 -07:00
$Y.CORP.YAHOO.COM\tgraves$ Y.CORP.YAHOO.COM\tgraves	9d4246863a	Support distributed cache files and archives on spark on yarn and attempt to cleanup the staging directory on exit	2013-09-23 09:09:59 -05:00
Nick Pentreath	d952f04c8e	Merge remote-tracking branch 'upstream/master' into implicit-als	2013-09-23 13:07:40 +02:00
Kay Ousterhout	c75eb14fe5	Send Task results through the block manager when larger than Akka frame size. This change requires adding an extra failure mode: tasks can complete successfully, but the result gets lost or flushed from the block manager before it's been fetched.	2013-09-22 21:20:48 -07:00
Holden Karau	7fe0b0ff56	Switch indent from 2 to 4 spaces	2013-09-22 19:44:51 -07:00
Reynold Xin	834686b108	Merge pull request #928 from jerryshao/fairscheduler-refactor Refactor FairSchedulableBuilder	2013-09-22 15:06:48 -07:00
Harvey	ef34cfb26c	Move Configuration broadcasts to SparkContext.	2013-09-22 14:43:58 -07:00
Harvey	a6eeb5ffd5	Add a cache for HadoopRDD metadata needed during computation. Currently, the cache is in SparkHadoopUtils, since it's conveniently a member of the SparkEnv.	2013-09-22 03:09:17 -07:00
jerryshao	77e9da1f34	Change Exception to NoSuchElementException and minor style fix	2013-09-22 16:50:08 +08:00
jerryshao	85024acd2e	Remove infix style and others	2013-09-22 14:20:55 +08:00
jerryshao	5850f599dd	Refactor FairSchedulableBuilder: 1. Configuration can be read from classpath if not set explicitly. 2. Add missing close handler.	2013-09-22 14:20:55 +08:00
Reynold Xin	a2ea069a5f	Merge pull request #937 from jerryshao/localProperties-fix Fix PR926 local properties issues in Spark Streaming like scenarios	2013-09-21 23:04:42 -07:00
Reynold Xin	f06f2da2cb	Merge pull request #941 from ilikerps/master Add "org.apache." prefix to packages in spark-class	2013-09-21 22:43:34 -07:00
Reynold Xin	7bb12a2af3	Merge pull request #940 from ankurdave/clear-port-properties-after-tests After unit tests, clear port properties unconditionally	2013-09-21 22:42:46 -07:00
Harvey	be0fc7246f	Split HadoopRDD into one for general Hadoop datasets and one tailored to Hadoop files, which is a common case. This is the first step to avoiding unnecessary Configuration broadcasts per HadoopRDD instantiation.	2013-09-21 21:14:14 -07:00
jerryshao	aa0c29f747	Add barrier for local properties unit test and fix some styles	2013-09-22 09:53:11 +08:00
Aaron Davidson	8933f9e98e	Add "org.apache." prefix to packages in spark-class Lacking this, the if/case statements never trigger on Spark 0.8.0+.	2013-09-20 19:27:08 -07:00
Reynold Xin	42571d30d0	Smarter take/limit implementation.	2013-09-20 17:09:53 -07:00
Reynold Xin	119de80294	Merge branch 'master' of github.com:mesos/spark	2013-09-20 15:03:55 -07:00
Vadim Chekan	fbe40c5806	Serialize and restore spark.cleaner.ttl to savepoint	2013-09-20 12:13:48 -07:00
Reynold Xin	1d87616b61	Made output of CoGroup and aggregations interruptible.	2013-09-19 23:31:36 -07:00
Mike	9524b943a4	Synchronize on "entries" the remaining update to "currentMemory". Make "currentMemory" @volatile, so that it's reads in ensureFreeSpace() are atomic and up-to-date--i.e., currentMemory can't increase while putLock is held (though it could decrease, which would only help ensureFreeSpace()).	2013-09-19 23:31:35 -07:00
Ankur Dave	026dba6aba	After unit tests, clear port properties unconditionally In MapOutputTrackerSuite, the "remote fetch" test sets spark.driver.port and spark.hostPort, assuming that they will be cleared by LocalSparkContext. However, the test never sets sc, so it remains null, causing LocalSparkContext to skip clearing these properties. Subsequent tests therefore fail with java.net.BindException: "Address already in use". This commit makes LocalSparkContext clear the properties even if sc is null.	2013-09-19 22:05:23 -07:00
Reynold Xin	c5e40954eb	Wrap around cached data to InterruptibleIterator.	2013-09-19 18:44:38 -07:00
Reynold Xin	c68e72be59	Added comment to InterruptibleIterator.	2013-09-19 18:40:06 -07:00
Reynold Xin	70953810b4	Added task killing iterator to RDDs that take inputs.	2013-09-19 18:33:16 -07:00

... 3 4 5 6 7 ...

4329 commits