spark

by apache · Data

Apache Spark - A unified analytics engine for large-scale data processing

New0 ratings43,956 starsActive
DataDeveloper ToolScalaPython
Open project ↗GitHub⚑ Report
spark preview

About this project

Apache Spark Spark is a unified analytics engine for large-scale data processing. It provides high-level APIs in Scala, Java, Python, and R (Deprecated), and an optimized engine that supports general computation graphs for data analysis. It also supports a rich set of higher-level tools including Spark SQL for SQL and DataFrames, pandas API on Spark for pandas workloads, MLlib for machine learning, GraphX for graph processing, and Structured Streaming for stream processing. — Official version: — Development version: Online Documentation You can find the latest Spark documentation, including a programming guide, on the project web page. This README file only contains basic…

Technologies

PythonJavaJupyter NotebookScalabig-dataHiveQLjdbc

Project health

Actively maintained
Last update2 days ago
Contributors331
Latest release
Open issues & PRs503
LicenseApache-2.0
On GitHubsince 2014

GitHub

43,956
stars
29,370
forks
331
contributors
503
open issues & PRs
Scala
language
2 days ago
last commit
View on GitHub ↗

Reviews

out of 5 · 0 ratings
★★★★★
0%
★★★★
0%
★★★
0%
★★
0%
0%
Sign in to write a review
No reviews yet
Be the first to review spark.

Built by

apache
Imported from GitHub · not yet claimed on GitPalace
View developer pageSign in with GitHub to claim

Maintain apache/spark? Claiming verifies admin access through your GitHub account and gives you control of this listing.

You might also like