Showing 21 of 93 projects
A Spark library for reading from and writing to Google BigQuery using DataFrames and SQL.
A collection of libraries for large-scale data processing in Hadoop ecosystems, including Spark, Pig, and incremental MapReduce.
An experimental Rust client for Apache Spark Connect, providing a DataFrame API to interact with Spark clusters.
An open-source framework for developing large-scale anomaly detection models using Apache Spark.
A PMML evaluator library for Apache Spark that provides ML-compatible transformers for deploying predictive models.
A Docker container providing a complete streaming environment for experimenting with Kafka, Spark Streaming, and Cassandra.
A collection of interactive Jupyter notebooks for learning Hadoop, Spark, and MapReduce with hands-on tutorials and demos.
A Spark application for migrating data to ScyllaDB from CQL-compatible databases or DynamoDB via Alternator.
Demo code for analyzing AWS CloudTrail and S3 logs using Apache Spark to detect security anomalies and enable SQL queries.
A Scala sample application demonstrating Spark Streaming integration with Kafka and Cassandra for data processing.
A Scala/Spark library for efficient processing, extraction, and derivation of web archive data (CDX/WARC).
An open-source toolkit for analyzing line-oriented JSON Twitter archives using Apache Spark.
A splitable Hadoop InputFormat for processing concatenated GZIP files and web archive (*.warc.gz) data efficiently in distributed systems.
Automates AWS EMR cluster deployment and Spark job submission with integrated logging and PySparkling support for H2O.
A Gatling extension for stress testing SQL databases and Spark Thrift Server via JDBC.
A thin C# gRPC client for communicating with Apache Spark Connect servers, enabling .NET applications to interact with Spark clusters.
A sample Spark job demonstrating how to use Spark Jobserver to run Apache Spark analytics with Cassandra.
An Apache Spark plugin that collects system resource metrics (CPU, memory) not provided by Spark's native metrics system.
A sample Spark job demonstrating analytics on Cassandra with SSL encryption for secure data processing.
Import CSV files from AWS S3 into Cassandra using Apache Spark with a simple configuration-based approach.
A Java project that reads from and writes to TDengine using Apache Spark for data processing.
Open-Awesome is built by the community, for the community. Submit a project, suggest an awesome list, or help improve the catalog on GitHub.