Showing 36 of 258 projects
A collection of R packages for interacting with Hadoop ecosystems, enabling big data analysis from R.
A pure Go client library for interacting with HBase databases, supporting HBase >= 1.0.
A Docker image for Apache Spark on YARN, built on Hadoop and CentOS for easy deployment.
A lightweight real-time big data streaming engine built on Akka for high-throughput, low-latency data processing.
A Jupyter Notebook kernel for interactive data exploration and analysis using Apache Spark with Scala.
TensorFlow binding for Apache Spark DataFrames, enabling TensorFlow program execution on Spark data.
A low-code visual tool for domain experts to build, run, and monitor real-time decision algorithms on streaming data.
Official connector for integrating Apache Spark with MongoDB, enabling distributed data processing on MongoDB data.
A quick reference guide to the most commonly used patterns and functions in PySpark SQL.
A comprehensive learning guide and interview refresher for Apache Spark, covering core concepts, architecture, and performance optimization.
A PySpark library providing helper methods for DataFrame validation, column transformations, and schema utilities to boost developer productivity.
A large-scale data warehouse system that provides approximate query answers with error bounds on massive datasets up to 300x faster than Hive.
A high-performance C++/DPC++ library for accelerated machine learning on CPUs, GPUs, and distributed systems.
A command-line tool for launching Apache Spark clusters on AWS EC2 with fast, configurable deployments.
A unified resource scheduler for co-scheduling batch, stateless, and stateful workloads in a single cluster to maximize resource utilization.
The fastest delimited file reader for R, using lazy loading and multi-threading to achieve speeds over 1 GB/sec.
An R package providing a lightweight frontend to use Apache Spark for distributed data processing from R.
Mirror of Apache Giraph
A collection of sample bootstrap action scripts for configuring applications on Amazon EMR clusters.
A fully asynchronous, non-blocking, thread-safe, high-performance Java client for HBase.
A Clojure DSL for Apache Spark that enables distributed data processing using idiomatic Clojure.
A server-side secondary index implementation for Apache HBase 0.94.8 using co-processors to enable efficient indexed queries.
An open-source security analytics platform that integrates big data technologies for centralized security monitoring, threat detection, and investigation.
An optimized distributed gradient boosting library for fast and accurate machine learning on large datasets.
A high-performance, disk-backed queue library using memory-mapped files for fast, persistent, and thread-safe data processing.
A Clojure library for writing map-reduce queries that compile to Apache Pig or Cascading, enabling distributed data processing with Clojure syntax.
A Scala-based event data simulator that generates realistic web traffic for a fake music streaming service.
A library enabling Apache Spark to read from and write to Apache HBase tables as external data sources using DataFrames and SQL.
A .NET stream processing library for Apache Kafka, providing a Kafka Streams-like API for building real-time applications.
A collection of GIS tools for spatial analysis of big data using Hadoop, integrating with ArcGIS Geoprocessing.
A library for parsing and querying XML data with Apache Spark SQL and DataFrames.
A Spark Streaming library for mining big data streams with incremental learning algorithms.
Kotlin bindings and extensions for Apache Spark, enabling idiomatic Kotlin development with data classes, lambdas, and null safety.
A comprehensive suite of Java NLP libraries and tools for text annotation, feature extraction, and language processing tasks.
A visualization framework for Apache Pig workflows that combines graphical depictions with real-time execution information.
A fast Apache Spark testing helper library with beautifully formatted error messages for Scala applications.
Open-Awesome is built by the community, for the community. Submit a project, suggest an awesome list, or help improve the catalog on GitHub.