Apache Spark - A unified analytics engine for large-scale data processing

Go to file

Matei Zaharia 4b1646a25f Support for non-filesystem-based Hadoop data sources		2011-07-06 20:37:55 -04:00
bagel/src	Fixed unit tests by making them clean up the SparkContext after use and	2011-05-13 12:03:58 -07:00
conf	Removed java-opts.template	2011-05-24 15:59:01 -07:00
core	Support for non-filesystem-based Hadoop data sources	2011-07-06 20:37:55 -04:00
examples/src/main/scala/spark/examples	Merge branch 'mos-bt'	2011-06-26 18:22:12 -07:00
project	Move managedStyle to SparkProject.	2011-06-02 14:06:54 +01:00
repl	Print version number 0.3 in REPL	2011-06-26 18:27:01 -07:00
sbt	Give SBT a bit more memory so it can do a update / compile / test in one JVM	2011-05-31 23:33:47 -07:00
.gitignore	Added more SBT stuff to gitignore	2011-02-08 17:06:07 -08:00
LICENSE	Added BSD license	2010-12-07 10:32:17 -08:00
lr_data.txt	Initial commit	2010-03-29 16:17:55 -07:00
README.md	typo	2011-06-23 02:28:17 +02:00
run	Pass quoted arguments properly to run	2011-05-31 23:34:16 -07:00
spark-executor	Made spark-executor output slightly nicer	2010-09-29 00:22:09 -07:00
spark-shell	Initial commit	2010-03-29 16:17:55 -07:00

README.md

Spark

Lightning-Fast Cluster Computing - http://www.spark-project.org/

Online Documentation

You can find the latest Spark documentation, including a programming guide, on the project wiki at http://github.com/mesos/spark/wiki. This file only contains basic setup instructions.

Building

Spark requires Scala 2.8. This version has been tested with 2.8.1.final. Experimental support for Scala 2.9 is available in the scala-2.9 branch.

The project is built using Simple Build Tool (SBT), which is packaged with it. To build Spark and its example programs, run:

sbt/sbt update compile

To run Spark, you will need to have Scala's bin in your PATH, or you will need to set the SCALA_HOME environment variable to point to where you've installed Scala. Scala must be accessible through one of these methods on Mesos slave nodes as well as on the master.

To run one of the examples, use ./run <class> <params>. For example:

./run spark.examples.SparkLR local[2]

will run the Logistic Regression example locally on 2 CPUs.

Each of the example programs prints usage help if no params are given.

All of the Spark samples take a <host> parameter that is the Mesos master to connect to. This can be a Mesos URL, or "local" to run locally with one thread, or "local[N]" to run locally with N threads.

Configuration

Spark can be configured through two files: conf/java-opts and conf/spark-env.sh.

In java-opts, you can add flags to be passed to the JVM when running Spark.

In spark-env.sh, you can set any environment variables you wish to be available when running Spark programs, such as PATH, SCALA_HOME, etc. There are also several Spark-specific variables you can set:

SPARK_CLASSPATH: Extra entries to be added to the classpath, separated by ":".
SPARK_MEM: Memory for Spark to use, in the format used by java's -Xmx option (for example, -Xmx200m means 200 MB, -Xmx1g means 1 GB, etc).
SPARK_LIBRARY_PATH: Extra entries to add to java.library.path for locating shared libraries.
SPARK_JAVA_OPTS: Extra options to pass to JVM.

Note that spark-env.sh must be a shell script (it must be executable and start with a #! header to specify the shell to use).