ODIn/spark-instrumented-optimizer

Author	SHA1	Message	Date
Imran Rashid	8a0a5ed533	track total partitions, in addition to cached partitions; use scala string formatting	2013-02-01 00:23:38 -08:00
Imran Rashid	f127f2ae76	fixup merge (master -> driver renaming)	2013-02-01 00:20:49 -08:00
Reynold Xin	f9af9cee6f	Moved PruneDependency into PartitionPruningRDD.scala.	2013-02-01 00:02:46 -08:00
Matei Zaharia	7e2e046e37	Merge pull request #434 from pwendell/python-exceptions SPARK-673: Capture and re-throw Python exceptions	2013-01-31 21:58:26 -08:00
Patrick Wendell	39ab83e957	Small fix from last commit	2013-01-31 21:52:52 -08:00
Patrick Wendell	c33f0ef41a	Some style cleanup	2013-01-31 21:50:02 -08:00
Patrick Wendell	3446d5c8d6	SPARK-673: Capture and re-throw Python exceptions This patch alters the Python <-> executor protocol to pass on exception data when they occur in user Python code.	2013-01-31 18:06:11 -08:00
Reynold Xin	6289d9654e	Removed the TODO comment from PartitionPruningRDD.	2013-01-31 17:49:36 -08:00
Reynold Xin	5b0fc265c2	Changed PartitionPruningRDD's split to make sure it returns the correct split index.	2013-01-31 17:48:39 -08:00
Stephen Haberman	782187c210	Once we find a split with no block, we don't have to look for more.	2013-01-31 18:27:25 -06:00
Stephen Haberman	418e36caa8	Add more private declarations.	2013-01-31 17:18:33 -06:00
Mikhail Bautin	fe3eceab57	Remove activation of profiles by default See the discussion at https://github.com/mesos/spark/pull/355 for why default profile activation is a problem.	2013-01-31 13:30:41 -08:00
Imran Rashid	02a6761589	Merge branch 'master' into blockmanager_info Conflicts: core/src/main/scala/spark/storage/BlockManagerMaster.scala	2013-01-30 18:52:35 -08:00
Imran Rashid	c1df24d085	rename Slaves --> Executor	2013-01-30 18:51:14 -08:00
Matei Zaharia	d12330bd2c	Merge pull request #426 from woggling/conn-manager-ips Remember ConnectionManagerId used to initiate SendingConnections	2013-01-30 15:02:53 -08:00
Matei Zaharia	612a9fee71	Merge pull request #428 from woggling/mesos-exec-id Make ExecutorIDs include SlaveIDs when running Mesos	2013-01-30 15:01:46 -08:00
Stephen Haberman	871476d506	Include message and exitStatus if availalbe.	2013-01-30 16:56:46 -06:00
Charles Reiss	252845d304	Remove remants of attempt to use slaveId-executorId in MesosExecutorBackend	2013-01-30 10:38:06 -08:00
Charles Reiss	f7de6978c1	Use Mesos ExecutorIDs to hold SlaveIDs. Then we can safely use the Mesos ExecutorID as a Spark ExecutorID.	2013-01-30 09:38:57 -08:00
Charles Reiss	7f51458774	Comment at top of DAGSchedulerSuite	2013-01-30 09:34:53 -08:00
Charles Reiss	9c0bae75ad	Change DAGSchedulerSuite to run DAGScheduler in the same Thread.	2013-01-30 09:22:07 -08:00
Charles Reiss	178b89204c	Refactor DAGScheduler more to allow testing without a separate thread.	2013-01-30 09:19:55 -08:00
Charles Reiss	4bf3d7ea12	Clear spark.master.port to cleanup for other tests	2013-01-29 19:05:58 -08:00
Charles Reiss	9eac7d01f0	Add DAGScheduler tests.	2013-01-29 18:55:43 -08:00
Charles Reiss	a3d14c0404	Refactoring to DAGScheduler to aid testing	2013-01-29 18:55:42 -08:00
Charles Reiss	16a0789e10	Remember ConnectionManagerId used to initiate SendingConnections. This prevents ConnectionManager from getting confused if a machine has multiple host names and the one getHostName() finds happens not to be the one that was passed from, e.g., the BlockManagerMaster.	2013-01-29 18:13:59 -08:00
Matei Zaharia	d54b10b6ad	Merge remote-tracking branch 'stephenh/removefailedjob' Conflicts: core/src/main/scala/spark/deploy/master/Master.scala	2013-01-29 18:12:29 -08:00
Matei Zaharia	ccb67ff2ca	Merge pull request #425 from stephenh/toDebugString Add RDD.toDebugString.	2013-01-29 10:44:18 -08:00
Matei Zaharia	9ae11603b4	Merge pull request #415 from stephenh/driver Replace old 'master' term with 'driver'.	2013-01-29 10:41:42 -08:00
Charles Reiss	a34096a76d	Add easymock to POMs	2013-01-29 10:04:33 -08:00
Imran Rashid	b92259ba57	Merge branch 'master' into blockmanager_info	2013-01-29 09:45:10 -08:00
Matei Zaharia	64ba6a8c2c	Simplify checkpointing code and RDD class a little: - RDD's getDependencies and getSplits methods are now guaranteed to be called only once, so subclasses can safely do computation in there without worrying about caching the results. - The management of a "splits_" variable that is cleared out when we checkpoint an RDD is now done in the RDD class. - A few of the RDD subclasses are simpler. - CheckpointRDD's compute() method no longer assumes that it is given a CheckpointRDDSplit -- it can work just as well on a split from the original RDD, because it only looks at its index. This is important because things like UnionRDD and ZippedRDD remember the parent's splits as part of their own and wouldn't work on checkpointed parents. - RDD.iterator can now reuse cached data if an RDD is computed before it is checkpointed. It seems like it wouldn't do this before (it always called iterator() on the CheckpointRDD, which read from HDFS).	2013-01-28 22:30:12 -08:00
Stephen Haberman	cbf72bffa5	Include name, if set, in RDD.toString().	2013-01-29 00:20:36 -06:00
Stephen Haberman	3cda14af3f	Add number of splits.	2013-01-29 00:12:31 -06:00
Matei Zaharia	a1ecec8d79	Merge branch 'master' of github.com:mesos/spark	2013-01-28 22:08:44 -08:00
Stephen Haberman	951cfd9ba2	Add JavaRDDLike.toDebugString().	2013-01-29 00:02:17 -06:00
Matei Zaharia	f6eb1f0825	Merge pull request #413 from pwendell/stage-logging SPARK-658: Adding logging of stage duration	2013-01-28 22:01:52 -08:00
Stephen Haberman	b45857c965	Add RDD.toDebugString. Original idea by Nathan Kronenfeld.	2013-01-28 23:56:56 -06:00
Patrick Wendell	7ee824e42e	Units from ms -> s	2013-01-28 21:48:32 -08:00
Stephen Haberman	13368818af	Merge branch 'master' into driver Conflicts: core/src/main/scala/spark/SparkContext.scala core/src/main/scala/spark/SparkEnv.scala core/src/main/scala/spark/deploy/LocalSparkCluster.scala core/src/main/scala/spark/executor/StandaloneExecutorBackend.scala core/src/main/scala/spark/scheduler/cluster/SparkDeploySchedulerBackend.scala core/src/main/scala/spark/scheduler/cluster/StandaloneClusterMessage.scala core/src/main/scala/spark/scheduler/cluster/StandaloneSchedulerBackend.scala core/src/main/scala/spark/storage/BlockManagerMaster.scala core/src/main/scala/spark/storage/ThreadingTest.scala core/src/test/scala/spark/MapOutputTrackerSuite.scala	2013-01-28 23:30:24 -06:00
Matei Zaharia	dda2ce017c	Merge pull request #424 from pwendell/logging-cleanup Some DEBUG-level log cleanup.	2013-01-28 21:18:54 -08:00
Patrick Wendell	1f9b486a8b	Some DEBUG-level log cleanup. A few changes to make the DEBUG-level logs less noisy and more readable. - Moved a few very frequent messages to Trace - Changed some BlockManger log messages to make them more understandable SPARK-666 #resolve	2013-01-28 20:29:35 -08:00
Imran Rashid	efff7bfb33	add long and float accumulatorparams	2013-01-28 20:23:11 -08:00
Imran Rashid	0f22c4207f	better formatting for RDDInfo	2013-01-28 20:07:53 -08:00
Imran Rashid	a423ee546c	expose RDD & storage info directly via SparkContext	2013-01-28 20:07:53 -08:00
Patrick Wendell	501433f1d5	Making submission time a field	2013-01-28 10:45:57 -08:00
Patrick Wendell	c423be7d8e	Renaming stage finished function	2013-01-28 10:45:57 -08:00
Patrick Wendell	07f568e1bf	SPARK-658: Adding logging of stage duration	2013-01-28 10:45:57 -08:00
Matei Zaharia	286f8f876f	Change time unit in MetadataCleaner to seconds	2013-01-28 01:29:27 -08:00
Matei Zaharia	f03d9760fd	Clean up BlockManagerUI a little (make it not be an object, merge with Directives, and bind to a random port)	2013-01-27 23:56:14 -08:00
Matei Zaharia	909850729e	Rename more things from slave to executor	2013-01-27 23:17:20 -08:00
Matei Zaharia	44b4a0f88f	Track workers by executor ID instead of hostname to allow multiple executors per machine and remove the need for multiple IP addresses in unit tests.	2013-01-27 19:23:49 -08:00
Matei Zaharia	6ad8540b40	Merge pull request #401 from squito/blockmanager_ui Blockmanager ui	2013-01-27 15:51:08 -08:00
Matei Zaharia	49f6472c0f	Merge pull request #418 from woggling/reregister-deadlock Fix BlockManager reregistration deadlock; do BlockManager reregistration more asynchronously	2013-01-26 18:59:02 -08:00
Charles Reiss	58fc6b2bed	Handle duplicate registrations better.	2013-01-26 18:30:44 -08:00
Charles Reiss	ad4232b4da	Fix deadlock in BlockManager reregistration triggered by failed updates.	2013-01-26 18:30:38 -08:00
Josh Rosen	d49cf0e587	Fix JavaRDDLike.flatMap(PairFlatMapFunction) (SPARK-668). This workaround is easier than rewriting JavaRDDLike in Java.	2013-01-26 16:13:18 -08:00
Imran Rashid	49c05608f5	add metadatacleaner for persisentRdd map	2013-01-25 17:04:16 -08:00
Stephen Haberman	8efbda0b17	Call executeOnCompleteCallbacks in more finally blocks.	2013-01-25 14:55:33 -06:00
Imran Rashid	a1d9d1767d	fixup `1cadaa1`, changed api of map	2013-01-25 10:05:26 -08:00
Imran Rashid	1cadaa164e	switch to TimeStampedHashMap for storing persistent Rdds	2013-01-25 09:30:21 -08:00
Imran Rashid	539491bbc3	code reformatting	2013-01-25 09:29:59 -08:00
Stephen Haberman	7dfb82a992	Replace old 'master' term with 'driver'.	2013-01-25 11:03:00 -06:00
Stephen Haberman	ec43a51b38	Merge branch 'master' into localsparkcontext Conflicts: core/src/test/scala/spark/FileServerSuite.scala core/src/test/scala/spark/RDDSuite.scala	2013-01-24 21:17:30 -06:00
Patrick Wendell	b6fc6e6752	SPARK-541: Adding a warning for invalid Master URL Right now Spark silently parses master URL's which do not match any known regex as a Mesos URL. The Mesos error message when an invalid URL gets passed is really confusing, so this warns the user when the implicit conversion is happening.	2013-01-24 14:31:23 -08:00
Stephen Haberman	230bda2047	Add LocalSparkContext to manage common sc variable.	2013-01-24 11:01:01 -06:00
Matei Zaharia	0fe173a3a5	Merge pull request #410 from rxin/splitpruningrdd Added a clearDependencies method in PartitionPruningRDD.	2013-01-23 23:10:15 -08:00
Reynold Xin	67a43bc7e6	Added a clearDependencies method in PartitionPruningRDD.	2013-01-23 23:06:52 -08:00
Matei Zaharia	fe5e4812fc	Merge pull request #409 from rxin/splitpruningrdd Added pruntSplits method to RDD.	2013-01-23 22:23:22 -08:00
Reynold Xin	c109f29c97	Updated PruneDependency to change "split" to "partition".	2013-01-23 22:22:03 -08:00
Reynold Xin	eedc542a02	Removed pruneSplits method in RDD and renamed SplitsPruningRDD to PartitionPruningRDD.	2013-01-23 22:14:23 -08:00
Reynold Xin	81004b967e	Marked prev RDD as transient in SplitsPruningRDD.	2013-01-23 21:54:27 -08:00
Reynold Xin	636e912f32	Created a PruneDependency to properly assign dependency for SplitsPruningRDD.	2013-01-23 21:21:55 -08:00
Reynold Xin	45cd50d5fe	Updated assert == to ===.	2013-01-23 16:06:58 -08:00
Matei Zaharia	548856a224	Merge remote-tracking branch 'woggling/remove-machines' Conflicts: core/src/main/scala/spark/scheduler/DAGScheduler.scala	2013-01-23 15:44:17 -08:00
Reynold Xin	c24b3819dd	Added an extra assert for split size check.	2013-01-23 15:34:59 -08:00
Reynold Xin	eb222b7206	Added pruntSplits method to RDD.	2013-01-23 15:29:02 -08:00
Matei Zaharia	1dd82743e0	Fix compile error due to cherry-pick	2013-01-23 13:07:27 -08:00
Charles Reiss	5c7422292e	Remove more dead code from test.	2013-01-23 12:59:51 -08:00
Imran Rashid	e1985bfa04	be sure to set class loader of kryo instances	2013-01-23 12:51:09 -08:00
Charles Reiss	be4a115a7e	Clarify TODO.	2013-01-23 12:48:45 -08:00
Charles Reiss	88b9d240fd	Remove dead code in test.	2013-01-23 12:40:38 -08:00
Matei Zaharia	1a3aeeca23	Merge pull request #407 from woggling/no-cache-tracker Eliminate CacheTracker	2013-01-23 12:28:48 -08:00
Charles Reiss	e1027ca639	Actually add CacheManager.	2013-01-23 12:22:11 -08:00
Matei Zaharia	4147e1d47b	Merge pull request #406 from tdas/master Changed StorageLevel and BlockManagerId API to prevent duplication in memory	2013-01-23 12:18:31 -08:00
Matei Zaharia	4d77d554e1	Merge pull request #394 from JoshRosen/add_file_fix Add SparkFiles.get() API to access files added through addFile().	2013-01-23 12:16:30 -08:00
Josh Rosen	ae2ed2947d	Allow PySpark's SparkFiles to be used from driver Fix minor documentation formatting issues.	2013-01-23 10:58:50 -08:00
Tathagata Das	79d55700ce	One more fix. Made even default constructor of BlockManagerId private to prevent such problems in the future.	2013-01-23 01:57:09 -08:00
Charles Reiss	0b506dd2ec	Add tests of various node failure scenarios.	2013-01-23 01:38:15 -08:00
Charles Reiss	d209b6b764	Extra debugging from hostLost()	2013-01-23 01:35:14 -08:00
Charles Reiss	9a27062260	Force generation increment after shuffle map stage	2013-01-23 01:34:44 -08:00
Tathagata Das	155f31398d	Made StorageLevel constructor private, and added StorageLevels.create() to the Java API. Updates scala and java programming guides.	2013-01-23 01:10:26 -08:00
Tathagata Das	5e11f1e51f	Modified StorageLevel API to ensure zero duplicate objects.	2013-01-22 23:42:53 -08:00
Tathagata Das	bacade6caf	Modified BlockManagerId API to ensure zero duplicate objects. Fixed BlockManagerId testcase in BlockManagerTestSuite.	2013-01-22 22:55:26 -08:00
Josh Rosen	43e9ff9596	Add test for driver hanging on exit (SPARK-530).	2013-01-22 22:47:26 -08:00
Charles Reiss	2849931000	Eliminate CacheTracker. Replaces DAGScheduler's queries of CacheTracker with BlockManagerMaster queries. Adds CacheManager to locally coordinate computation of cached RDDs.	2013-01-22 22:19:30 -08:00
Matei Zaharia	ebaa8f6519	Merge remote-tracking branch 'stephenh/cleanup' Conflicts: core/src/main/scala/spark/scheduler/local/LocalScheduler.scala	2013-01-22 21:05:45 -08:00
Matei Zaharia	d2d273868b	Merge pull request #397 from JoshRosen/refactoring/daemon-threads Refactor daemon thread creation	2013-01-22 21:02:53 -08:00
Stephen Haberman	98d0b7747d	Fix Worker logInfo about unknown executor.	2013-01-22 18:11:51 -06:00
Stephen Haberman	8c51322cd0	Don't bother creating an exception.	2013-01-22 18:09:10 -06:00

1 2 3 4 5 ...

1066 commits