Showing 36 of 258 projects
Mozilla's utility library for Hadoop, HBase, Pig, and related big data technologies.
A PHP client extension for the TDengine big data engine, with Swoole coroutine support.
A simple utility for testing Apache Hive scripts locally without requiring Java development skills.
Go implementation of Count-Min-Log sketch for improved approximate counting of low-frequency events.
A unit test framework for Hive scripts that provides an embedded Hive environment with Derby database and HiveThriftService.
Code samples demonstrating how to use popular applications on Amazon Elastic MapReduce (EMR).
Jupyter notebooks for hands-on Big Data Analytics exercise classes covering Spark ML, Map/Reduce algorithms, and deep learning.
A Ruby gem for extracting data from JSON streams based on keys, nesting levels, or custom conditions without implementing low-level callbacks.
An Erlang/Elixir driver for HBase that uses a Java server with Asynchbase to provide asynchronous database queries.
A distributed platform for processing continuous unbounded streams of data with a clean API and fault tolerance.
An Elasticsearch river plugin for importing data from HBase into Elasticsearch.
A lightweight ETL library and data integration toolbox for .NET, enabling programmatic data flow pipelines.
A massively-parallel C++ SQL query engine for lightning-fast analytics on petabytes of data in Hadoop clusters.
Rust bindings and safe wrapper APIs for Hadoop's libhdfs, enabling HDFS access from Rust applications.
clusterdock is a framework for creating Docker-based container clusters
Example notebooks for analyzing web archives using the Archives Unleashed Toolkit.
A Delphi framework for connecting to and performing CRUD operations on Apache Cassandra databases using the DataStax C/C++ driver.
A fast, highly-scalable graph database supporting over 10 billion vertices and edges with OLTP capabilities and dual Gremlin/Cypher query language support.
A tool for testing the DataStax Spark Connector against Apache Cassandra or DataStax Enterprise.
An Elixir MapReduce framework with Hadoop Streaming integration, simplifying distributed data processing.
A Scala sample application demonstrating Spark Streaming integration with Kafka and Cassandra for data processing.
A big data cluster management tool that creates and manages multi-technology clusters with health monitoring.
A fast HyperLogLog implementation for Elixir/Erlang that counts unique values with minimal memory usage.
A Delphi class framework for interacting with Couchbase NoSQL databases, providing CRUD operations, JSON subdocument manipulation, and N1QL query support.
A simple wrapper around cascading.hbase for seamless HBase integration in Cascalog workflows.
A Scala/Spark library for efficient processing, extraction, and derivation of web archive data (CDX/WARC).
A distributed data stream pipeline for querying, augmenting, and transforming data using Elixir pattern-matching rules.
Helm charts and Docker Compose configurations for deploying TDengine time-series database clusters on Kubernetes and Docker.
A lightweight, HDFS-compatible file system built over Cassandra with a fat driver design for easy deployment.
Provides HBase adapters for reading and writing data within Cascading data processing workflows on Hadoop clusters.
An open-source toolkit for analyzing line-oriented JSON Twitter archives using Apache Spark.
A splitable Hadoop InputFormat for processing concatenated GZIP files and web archive (*.warc.gz) data efficiently in distributed systems.
Code for the Best Buy competition at Kaggle, focusing on mobile contest big data analysis.
Example code for large-scale metrics analysis using Ruby with AWS services like EMR and Redshift.
A thin C# gRPC client for communicating with Apache Spark Connect servers, enabling .NET applications to interact with Spark clusters.
A sample Spark job demonstrating how to use Spark Jobserver to run Apache Spark analytics with Cassandra.
Open-Awesome is built by the community, for the community. Submit a project, suggest an awesome list, or help improve the catalog on GitHub.