ODIn/spark-instrumented-optimizer

Author	SHA1	Message	Date
Reynold Xin	ac712e48af	Merge pull request #524 from rxin/doc Added spark.shuffle.file.buffer.kb to configuration doc. Author: Reynold Xin <rxin@apache.org> == Merge branch commits == commit 0eea1d761ff772ff89be234e1e28035d54e5a7de Author: Reynold Xin <rxin@apache.org> Date: Wed Jan 29 14:40:48 2014 -0800 Added spark.shuffle.file.buffer.kb to configuration doc.	2014-01-30 09:33:18 -08:00
Tathagata Das	7930209614	Merge pull request #497 from tdas/docs-update Updated Spark Streaming Programming Guide Here is the updated version of the Spark Streaming Programming Guide. This is still a work in progress, but the major changes are in place. So feedback is most welcome. In general, I have tried to make the guide to easier to understand even if the reader does not know much about Spark. The updated website is hosted here - http://www.eecs.berkeley.edu/~tdas/spark_docs/streaming-programming-guide.html The major changes are: - Overview illustrates the usecases of Spark Streaming - various input sources and various output sources - An example right after overview to quickly give an idea of what Spark Streaming program looks like - Made Java API and examples a first class citizen like Scala by using tabs to show both Scala and Java examples (similar to AMPCamp tutorial's code tabs) - Highlighted the DStream operations updateStateByKey and transform because of their powerful nature - Updated driver node failure recovery text to highlight automatic recovery in Spark standalone mode - Added information about linking and using the external input sources like Kafka and Flume - In general, reorganized the sections to better show the Basic section and the more advanced sections like Tuning and Recovery. Todos: - Links to the docs of external Kafka, Flume, etc - Illustrate window operation with figure as well as example. Author: Tathagata Das <tathagata.das1565@gmail.com> == Merge branch commits == commit 18ff10556570b39d672beeb0a32075215cfcc944 Author: Tathagata Das <tathagata.das1565@gmail.com> Date: Tue Jan 28 21:49:30 2014 -0800 Fixed a lot of broken links. commit 34a5a6008dac2e107624c7ff0db0824ee5bae45f Author: Tathagata Das <tathagata.das1565@gmail.com> Date: Tue Jan 28 18:02:28 2014 -0800 Updated github url to use SPARK_GITHUB_URL variable. commit f338a60ae8069e0a382d2cb170227e5757cc0b7a Author: Tathagata Das <tathagata.das1565@gmail.com> Date: Mon Jan 27 22:42:42 2014 -0800 More updates based on Patrick and Harvey's comments. commit 89a81ff25726bf6d26163e0dd938290a79582c0f Author: Tathagata Das <tathagata.das1565@gmail.com> Date: Mon Jan 27 13:08:34 2014 -0800 Updated docs based on Patricks PR comments. commit d5b6196b532b5746e019b959a79ea0cc013a8fc3 Author: Tathagata Das <tathagata.das1565@gmail.com> Date: Sun Jan 26 20:15:58 2014 -0800 Added spark.streaming.unpersist config and info on StreamingListener interface. commit e3dcb46ab83d7071f611d9b5008ba6bc16c9f951 Author: Tathagata Das <tathagata.das1565@gmail.com> Date: Sun Jan 26 18:41:12 2014 -0800 Fixed docs on StreamingContext.getOrCreate. commit 6c29524639463f11eec721e4d17a9d7159f2944b Author: Tathagata Das <tathagata.das1565@gmail.com> Date: Thu Jan 23 18:49:39 2014 -0800 Added example and figure for window operations, and links to Kafka and Flume API docs. commit f06b964a51bb3b21cde2ff8bdea7d9785f6ce3a9 Author: Tathagata Das <tathagata.das1565@gmail.com> Date: Wed Jan 22 22:49:12 2014 -0800 Fixed missing endhighlight tag in the MLlib guide. commit 036a7d46187ea3f2a0fb8349ef78f10d6c0b43a9 Merge: eab351d `a1cd185` Author: Tathagata Das <tathagata.das1565@gmail.com> Date: Wed Jan 22 22:17:42 2014 -0800 Merge remote-tracking branch 'apache/master' into docs-update commit eab351d05c0baef1d4b549e1581310087158d78d Author: Tathagata Das <tathagata.das1565@gmail.com> Date: Wed Jan 22 22:17:15 2014 -0800 Update Spark Streaming Programming Guide.	2014-01-28 21:51:05 -08:00
Reynold Xin	84670f2715	Merge pull request #466 from liyinan926/file-overwrite-new Allow files added through SparkContext.addFile() to be overwritten This is useful for the cases when a file needs to be refreshed and downloaded by the executors periodically. For example, a possible use case is: the driver periodically renews a Hadoop delegation token and writes it to a token file. The token file needs to be downloaded by the executors whenever it gets renewed. However, the current implementation throws an exception when the target file exists and its contents do not match those of the new source. This PR adds an option to allow files to be overwritten to support use cases similar to the above.	2014-01-27 17:08:35 -08:00
Josh Rosen	4cebb79c9f	Deprecate mapPartitionsWithSplit in PySpark. Also, replace the last reference to it in the docs. This fixes SPARK-1026.	2014-01-23 20:01:36 -08:00
Patrick Wendell	576c4a4c50	Merge pull request #478 from sryza/sandy-spark-1033 SPARK-1033. Ask for cores in Yarn container requests Tested on a pseudo-distributed cluster against the Fair Scheduler and observed a worker taking more than a single core.	2014-01-22 14:10:07 -08:00
Matei Zaharia	d009b17d13	Merge pull request #315 from rezazadeh/sparsesvd Sparse SVD # Singular Value Decomposition Given an m x n matrix A, compute matrices U, S, V such that A = U S * V^T* There is no restriction on m, but we require n^2 doubles to fit in memory. Further, n should be less than m. The decomposition is computed by first computing A^TA = V S^2 V^T, computing svd locally on that (since n x n is small), from which we recover S and V. Then we compute U via easy matrix multiplication as U = A V * S^-1* Only singular vectors associated with the largest k singular values If there are k such values, then the dimensions of the return will be: * S is k x k and diagonal, holding the singular values on diagonal. * U is m x k and satisfies U^TU = eye(k). V is n x k and satisfies V^TV = eye(k). All input and output is expected in sparse matrix format, 0-indexed as tuples of the form ((i,j),value) all in RDDs. # Testing Tests included. They test: - Decomposition promise (A = USV^T) - For small matrices, output is compared to that of jblas - Rank 1 matrix test included - Full Rank matrix test included - Middle-rank matrix forced via k included # Example Usage import org.apache.spark.SparkContext import org.apache.spark.mllib.linalg.SVD import org.apache.spark.mllib.linalg.SparseMatrix import org.apache.spark.mllib.linalg.MatrixyEntry // Load and parse the data file val data = sc.textFile("mllib/data/als/test.data").map { line => val parts = line.split(',') MatrixEntry(parts(0).toInt, parts(1).toInt, parts(2).toDouble) } val m = 4 val n = 4 // recover top 1 singular vector val decomposed = SVD.sparseSVD(SparseMatrix(data, m, n), 1) println("singular values = " + decomposed.S.data.toArray.mkString) # Documentation Added to docs/mllib-guide.md	2014-01-22 14:01:30 -08:00
Andrew Ash	069bb94206	Clarify spark.default.parallelism It's the task count across the cluster, not per worker, per machine, per core, or anything else.	2014-01-21 14:49:35 -08:00
Sandy Ryza	adf42611f1	Incorporate Tom's comments - update doc and code to reflect that core requests may not always be honored	2014-01-21 00:38:02 -08:00
Patrick Wendell	c324ac10ee	Force use of LZF when spilling data	2014-01-20 19:00:48 -08:00
Patrick Wendell	cdb003e376	Removing docs on akka options	2014-01-20 16:40:58 -08:00
Sandy Ryza	3e85b87d90	SPARK-1033. Ask for cores in Yarn container requests	2014-01-20 14:42:32 -08:00
Yinan Li	584323c6b1	Addressed comments from Reynold Signed-off-by: Yinan Li <liyinan926@gmail.com>	2014-01-18 21:28:17 -08:00
Patrick Wendell	bf5699543b	Merge pull request #462 from mateiz/conf-file-fix Remove Typesafe Config usage and conf files to fix nested property names With Typesafe Config we had the subtle problem of no longer allowing nested property names, which are used for a few of our properties: http://apache-spark-developers-list.1001551.n3.nabble.com/Config-properties-broken-in-master-td208.html This PR is for branch 0.9 but should be added into master too. (cherry picked from commit `34e911ce9a`) Signed-off-by: Patrick Wendell <pwendell@gmail.com>	2014-01-18 16:20:00 -08:00
Yinan Li	fd833e7ab1	Allow files added through SparkContext.addFile() to be overwritten This is useful for the cases when a file needs to be refreshed and downloaded by the executors periodically. Signed-off-by: Yinan Li <liyinan926@gmail.com>	2014-01-18 15:26:59 -08:00
Reza Zadeh	caf97a25a2	Merge remote-tracking branch 'upstream/master' into sparsesvd	2014-01-17 14:34:03 -08:00
Reza Zadeh	5c639d70df	0index docs	2014-01-17 14:31:39 -08:00
Reza Zadeh	cb13b15a60	use 0-indexing	2014-01-17 13:55:42 -08:00
Reza Zadeh	d28bf41827	changes from PR	2014-01-17 13:39:40 -08:00
Reynold Xin	0675ca50f3	Merge pull request #439 from CrazyJvm/master SPARK-1024 Remove "-XX:+UseCompressedStrings" option from tuning guide remove "-XX:+UseCompressedStrings" option from tuning guide since jdk7 no longer supports this.	2014-01-15 16:09:03 -08:00
Matei Zaharia	2ffdaefbcb	Clarify that Python 2.7 is only needed for MLlib	2014-01-15 14:20:39 -08:00
Patrick Wendell	494d3c0774	Merge pull request #433 from markhamstra/debFix Updated Debian packaging	2014-01-15 10:00:50 -08:00
CrazyJvm	263933da97	remove "-XX:+UseCompressedStrings" option remove "-XX:+UseCompressedStrings" option from tuning guide since jdk7 no longer supports this.	2014-01-15 22:26:15 +08:00
Reynold Xin	3d9e66d92a	Merge pull request #436 from ankurdave/VertexId-case Rename VertexID -> VertexId in GraphX	2014-01-14 23:17:05 -08:00
Mark Hamstra	147a943df0	Removed repl-bin and updated maven build doc.	2014-01-14 22:17:24 -08:00
Ankur Dave	f4d9019aa8	VertexID -> VertexId	2014-01-14 22:17:18 -08:00
Reynold Xin	3a386e2389	Merge pull request #424 from jegonzal/GraphXProgrammingGuide Additional edits for clarity in the graphx programming guide. Added an overview of the Graph and GraphOps functions and fixed numerous typos.	2014-01-14 21:52:50 -08:00
Ankur Dave	1210ec2945	Describe GraphX caching and uncaching in guide	2014-01-14 17:25:38 -08:00
Joseph E. Gonzalez	0bba7738a2	Additional edits for clarity in the graphx programming guide.	2014-01-14 10:31:54 -08:00
Joseph E. Gonzalez	486f37c59c	Improving the graphx-programming-guide.	2014-01-14 09:43:33 -08:00
Patrick Wendell	980250b1ee	Merge pull request #416 from tdas/filestream-fix Removed unnecessary DStream operations and updated docs Removed StreamingContext.registerInputStream and registerOutputStream - they were useless. InputDStream has been made to register itself, and just registering a DStream as output stream cause RDD objects to be created but the RDDs will not be computed at all.. Also made DStream.register() private[streaming] for the same reasons. Updated docs, specially added package documentation for streaming package. Also, changed NetworkWordCount's input storage level to use MEMORY_ONLY, replication on the local machine causes warning messages (as replication fails) which is scary for a new user trying out his/her first example.	2014-01-14 00:05:37 -08:00
Tathagata Das	f8bd828c7c	Fixed loose ends in docs.	2014-01-14 00:03:46 -08:00
Tathagata Das	f8e239e058	Merge remote-tracking branch 'apache/master' into filestream-fix Conflicts: streaming/src/main/scala/org/apache/spark/streaming/dstream/DStream.scala	2014-01-13 23:57:27 -08:00
Reza Zadeh	845e568fad	Merge remote-tracking branch 'upstream/master' into sparsesvd	2014-01-13 23:52:34 -08:00
Patrick Wendell	0984647aae	Enable compression by default for spills	2014-01-13 23:25:25 -08:00
Tathagata Das	4e497db8f3	Removed StreamingContext.registerInputStream and registerOutputStream - they were useless as InputDStream has been made to register itself. Also made DStream.register() private[streaming] - not useful to expose the confusing function. Updated a lot of documentation.	2014-01-13 23:23:46 -08:00
Patrick Wendell	fdaabdc673	Merge pull request #380 from mateiz/py-bayes Add Naive Bayes to Python MLlib, and some API fixes - Added a Python wrapper for Naive Bayes - Updated the Scala Naive Bayes to match the style of our other algorithms better and in particular make it easier to call from Java (added builder pattern, removed default value in train method) - Updated Python MLlib functions to not require a SparkContext; we can get that from the RDD the user gives - Added a toString method in LabeledPoint - Made the Python MLlib tests run as part of run-tests as well (before they could only be run individually through each file)	2014-01-13 23:08:26 -08:00
Patrick Wendell	4a805aff5e	Merge pull request #367 from ankurdave/graphx GraphX: Unifying Graphs and Tables GraphX extends Spark's distributed fault-tolerant collections API and interactive console with a new graph API which leverages recent advances in graph systems (e.g., [GraphLab](http://graphlab.org)) to enable users to easily and interactively build, transform, and reason about graph structured data at scale. See http://amplab.github.io/graphx/. Thanks to @jegonzal, @rxin, @ankurdave, @dcrankshaw, @jianpingjwang, @amatsukawa, @kellrott, and @adamnovak. Tasks left: - [x] Graph-level uncache - [x] Uncache previous iterations in Pregel - [x] ~~Uncache previous iterations in GraphLab~~ (postponed to post-release) - [x] - Describe GC issue with GraphLab - [ ] Write `docs/graphx-programming-guide.md` - [x] - Mention future Bagel support in docs - [ ] - Section on caching/uncaching in docs: As with Spark, cache something that is used more than once. In an iterative algorithm, try to cache and force (i.e., materialize) something every iteration, then uncache the cached things that depended on the newly materialized RDD but that won't be referenced again. - [x] Undo modifications to core collections and instead copy them to org.apache.spark.graphx - [x] Make Graph serializable to work around capture in Spark shell - [x] Rename graph -> graphx in package name and subproject - [x] Remove standalone PageRank - [x] ~~Fix amplab/graphx#52 by checking `iter.hasNext`~~	2014-01-13 22:58:38 -08:00
Patrick Wendell	945fe7a37e	Merge pull request #408 from pwendell/external-serializers Improvements to external sorting 1. Adds the option of compressing outputs. 2. Adds batching to the serialization to prevent OOM on the read side. 3. Slight renaming of config options. 4. Use Spark's buffer size for reads in addition to writes.	2014-01-13 22:56:12 -08:00
Joseph E. Gonzalez	4bafc4f41f	adding documentation about EdgeRDD	2014-01-13 22:55:54 -08:00
Ankur Dave	af645be5b8	Fix all code examples in guide	2014-01-13 22:29:45 -08:00
Ankur Dave	2cd9358ccf	Finish `6f6f8c928c`	2014-01-13 22:29:23 -08:00
Ankur Dave	6f6f8c928c	Wrap methods in the appropriate class/object declaration	2014-01-13 21:55:35 -08:00
Ankur Dave	67795dbbfb	Write Graph Builders section in guide	2014-01-13 21:45:11 -08:00
Ankur Dave	e14a14bcde	Remove K-Core and LDA sections from guide; they are unimplemented	2014-01-13 21:12:58 -08:00
Ankur Dave	59e4384e19	Fix Pregel SSSP example in programming guide	2014-01-13 21:02:38 -08:00
Joseph E. Gonzalez	ee8931d2c6	Finished documenting vertexrdd.	2014-01-13 19:30:35 -08:00
Joseph E. Gonzalez	552de5d42e	Finished second pass on pregel docs.	2014-01-13 18:40:43 -08:00
Joseph E. Gonzalez	622b7f7d39	Minor changes in graphx programming guide.	2014-01-13 18:40:43 -08:00
Joseph E. Gonzalez	cfe4a29dcb	Improvements in example code for the programming guide as well as adding serialization support for GraphImpl to address issues with failed closure capture.	2014-01-13 17:18:31 -08:00
Ankur Dave	1bd5cefcae	Remove aggregateNeighbors	2014-01-13 17:03:03 -08:00

1 2 3 4 5 ...

508 commits