ODIn/spark-instrumented-optimizer

Author	SHA1	Message	Date
Tathagata Das	3f08e1129f	Merge branch 'master' into td-rdd-save Conflicts: core/src/main/scala/spark/SparkContext.scala	2011-06-27 13:43:44 -07:00
Tathagata Das	ad842ac823	Merge branch 'master' into td-rdd-save Conflicts: core/src/main/scala/spark/RDD.scala	2011-06-27 13:39:11 -07:00
Matei Zaharia	c4dd68ae21	Merge branch 'mos-bt' This merge keeps only the broadcast work in mos-bt because the structure of shuffle has changed with the new RDD design. We still need some kind of parallel shuffle but that will be added later. Conflicts: core/src/main/scala/spark/BitTorrentBroadcast.scala core/src/main/scala/spark/ChainedBroadcast.scala core/src/main/scala/spark/RDD.scala core/src/main/scala/spark/SparkContext.scala core/src/main/scala/spark/Utils.scala core/src/main/scala/spark/shuffle/BasicLocalFileShuffle.scala core/src/main/scala/spark/shuffle/DfsShuffle.scala	2011-06-26 18:22:12 -07:00
Tathagata Das	38f2ba99cc	Further changes to HadoopFileWriter. Implemented ability to save RDDs as SequenceFiles and ObjectFiles. 1> HadoopFileWriter changed to take class types as constructor parameters (no more generic type) 2> Multiple types of RDD.saveAsHadoopFile() implemented to provide more saving options 3> RDD.saveAsSequenceFile() automatically converts basic types to Writable types before saving as SequenceFile 4> RDD.saveAsObjectFile() serializes objects and saves them to a ObjectFile 5> SparkContext.objectFile() opens the saved ObjectFiles	2011-06-24 19:51:21 -07:00
Olivier Grisel	2e3531d8bf	Implemented RDD.leftOuterJoin and RDD.rightOuterJoin	2011-06-24 11:00:51 +02:00
Tathagata Das	3d2befe831	Improved HadoopFileWriter (saves key and value classes to jobconf)	2011-06-23 08:11:22 -07:00
Matei Zaharia	214250016a	Added simple version of lookup	2011-06-20 11:59:16 -07:00
Matei Zaharia	23b1c309fb	Added pipe() operation on RDDs for mapping through a shell command.	2011-06-19 23:05:19 -07:00
Tathagata Das	b5e6645505	Cleaner reimplementation of HadoopFileWriter. Introduced TaskContext. 1> HadoopFileWriter works correctly with task failures 2> It can also take an user specified JobConf object for configuration settings 3> A Task can now get information like stage ID, split ID, and attempt ID using TaskContext class 4> Minor changes in SparkContext, DAGScheduler and subclasses to allow specification of TaskContext as a parameter	2011-06-16 20:57:57 -07:00
Tathagata Das	869836a2fa	Implemented TaskContext to hold contextual information (jobID, taskID, attemptID) of a task	2011-06-10 19:47:28 -07:00
Tathagata Das	389e56156f	HadoopFileWriter changed to use Hadoop's OutputCommitter	2011-06-09 15:29:22 -07:00
Tathagata Das	24d845833c	First-cut implementation of RDD.SaveAsText	2011-06-05 04:14:43 -07:00
Matei Zaharia	9bb448a151	Catch Throwable instead of Exception in LocalScheduler and Executor. Fixes #57 .	2011-06-01 11:45:47 -07:00
Matei Zaharia	850fe3274e	Make the runJob API public. Fixes #56 .	2011-06-01 11:38:44 -07:00
Matei Zaharia	5166d76843	Ensure logging is initialized before spawning any threads to fix issue #45	2011-05-31 23:47:32 -07:00
Matei Zaharia	8b0390d344	Instantiate NullWritable properly in HadoopFile	2011-05-30 23:54:14 -07:00
Matei Zaharia	c501cff924	Executor was looking for the wrong constructor for ExecutorClassLoader	2011-05-29 16:15:59 -07:00
Ismael Juma	1396678baa	Move REPL classes to separate module.	2011-05-27 11:22:50 +01:00
Ismael Juma	164ef4c751	Use explicit asInstanceOf instead of misleading unchecked pattern matching. Also enable -unchecked warnings in SBT build file.	2011-05-27 07:57:10 +01:00
Ismael Juma	89c8ea2bb2	Replace deprecated `-` and `--` with suggested filterNot (which is uglier).	2011-05-26 22:22:37 +01:00
Ismael Juma	94f05683bd	Replace deprecated `first` with `head`.	2011-05-26 22:13:41 +01:00
Ismael Juma	0b6a862b68	Use math instead of Math as the latter is deprecated.	2011-05-26 22:06:36 +01:00
Ismael Juma	1f27d94c48	Use Array.iterator instead of Iterator.fromArray as the latter is deprecated.	2011-05-26 22:04:42 +01:00
Ismael Juma	1993a8e556	Use += instead of + for mutable sequences as the latter is deprecated.	2011-05-26 21:59:48 +01:00
Matei Zaharia	cec427e777	Fixed a bug with preferred locations having changed meaning in new RDDs	2011-05-22 17:12:29 -07:00
Matei Zaharia	4c888b2933	Fix queue type for executor	2011-05-22 16:42:05 -07:00
Matei Zaharia	bea3a33012	doc tweak	2011-05-22 16:03:41 -07:00
Matei Zaharia	9bde5a54cb	class loader fix	2011-05-22 16:00:41 -07:00
Matei Zaharia	91c07a33d9	Various fixes to serialization	2011-05-21 22:50:08 -07:00
Matei Zaharia	f61b61c4ac	Merge branch 'master' into new-rdds	2011-05-21 21:25:58 -07:00
Matei Zaharia	24a1e7f838	Scheduler can now recover from lost map outputs	2011-05-20 00:19:53 -07:00
Matei Zaharia	82329b0b28	Updated scheduler to support running on just some partitions of final RDD	2011-05-19 12:47:09 -07:00
Matei Zaharia	328e51b693	Various minor fixes	2011-05-19 11:19:25 -07:00
Matei Zaharia	fd1d255821	Stop objectifying various trackers, caches, etc.	2011-05-17 12:41:13 -07:00
Matei Zaharia	4db50e26c7	Fixed unit tests by making them clean up the SparkContext after use and thus clean up the various singletons (RDDCache, MapOutputTracker, etc). This isn't perfect yet (ideally we shouldn't use singleton objects at all) but we can fix that later.	2011-05-13 12:03:58 -07:00
Matei Zaharia	aca8150c52	Ensure that AddedToCache messages make it home before tasks finish	2011-05-13 11:43:52 -07:00
Matei Zaharia	16c886a581	Optimization for count()	2011-05-13 10:41:34 -07:00
Mosharaf Chowdhury	db7a2c4897	Issue #42 fixed.	2011-04-28 14:30:48 -07:00
Ankur Dave	a4c04f3f6f	Error handling for disk I/O in DiskSpillingCache Also renamed the property spark.DiskSpillingCache.cacheDir to spark.diskSpillingCache.cacheDir in order to follow conventions.	2011-04-27 23:23:29 -07:00
Ankur Dave	12ff0d2dc3	Bring an entry back into memory after fetching it from disk	2011-04-27 22:59:05 -07:00
Ankur Dave	e30313aa2c	Added DiskSpillingCache DiskSpillingCache is a BoundedMemoryCache that spills entries to disk when it runs out of space. Currently the implementation is very simple. In particular, it's missing the following features: - Error handling for disk I/O, including checking of disk space levels - Bringing an entry back into memory after fetching it from disk In addition, here are some features that aren't critical but should be implemented soon: - Spilling based on a user-set priority in addition to LRU - Caching into a subdirectory of spark.DiskSpillingCache.cacheDir rather than the root directory	2011-04-27 22:32:35 -07:00
Mosharaf Chowdhury	60d1121343	Refactoring: daemonThreadFactories have all been moved to the Utils object instead of having multiple copies in Broadcast and Shuffle objects.	2011-04-27 22:13:01 -07:00
Mosharaf Chowdhury	e898e108a3	Cleanup + refactoring...	2011-04-27 22:00:24 -07:00
Mosharaf Chowdhury	0567646180	Shuffle is also working from its own subpackage.	2011-04-27 21:11:41 -07:00
Mosharaf Chowdhury	2742de707a	Removed some shuffle implementations. Remaining ones all use local files to write map outputs.	2011-04-27 20:53:43 -07:00
Mosharaf Chowdhury	9d78779257	Merge branch 'mos-shuffle-tracked' into mos-bt Conflicts: core/src/main/scala/spark/Broadcast.scala	2011-04-27 20:47:07 -07:00
Mosharaf Chowdhury	ac7e066383	Merge branch 'master' into mos-shuffle-tracked Conflicts: .gitignore core/src/main/scala/spark/LocalFileShuffle.scala src/scala/spark/BasicLocalFileShuffle.scala src/scala/spark/Broadcast.scala src/scala/spark/LocalFileShuffle.scala	2011-04-27 14:35:03 -07:00
Mosharaf Chowdhury	4e4c41026c	Added support for custom classes. (from 49ea48)	2011-04-27 12:30:16 -07:00
Mosharaf Chowdhury	65848da8df	Refacoring...	2011-04-26 17:41:31 -07:00
Mosharaf Chowdhury	b8ab7862b8	Moved broadcast-related code to separate directory under spark.broadcast package.	2011-04-26 17:22:52 -07:00

1 2 3

101 commits