Showing 22 of 58 projects
A Hadoop log aggregator and dashboard for visualizing cluster utilization across users.
A practical guide to exploratory data analytics using Hadoop with Pig and Ruby for terabyte-scale data processing.
An open-source toolkit for analyzing web archives at scale using Apache Spark.
A scalable malware processing and analytics platform built on Hadoop Pig for binary data extraction and analysis.
A Go database/sql driver for Apache Avatica server, enabling Go applications to connect to Phoenix and other Avatica-backed databases.
A collection of libraries for large-scale data processing in Hadoop ecosystems, including Spark, Pig, and incremental MapReduce.
An R extension for distributed computing using Apache Hive, enabling HQL queries in R and R functions in Hive.
A Scalding library for machine learning and statistical analysis, featuring Mahout vector integration, K-Means clustering, and Naive-Bayes classifiers.
A collection of interactive Jupyter notebooks for learning Hadoop, Spark, and MapReduce with hands-on tutorials and demos.
A production-grade HBase ORM library for clean, fast, and fun object-oriented data access, also compatible with Google Cloud Bigtable.
Mozilla's utility library for Hadoop, HBase, Pig, and related big data technologies.
A simple utility for testing Apache Hive scripts locally without requiring Java development skills.
A unit test framework for Hive scripts that provides an embedded Hive environment with Derby database and HiveThriftService.
Code samples demonstrating how to use popular applications on Amazon Elastic MapReduce (EMR).
MapReduce tools for bulk indexing of web archive WARC/ARC files into ZipNum sharded CDX clusters on Hadoop, EMR, or local systems.
A massively-parallel C++ SQL query engine for lightning-fast analytics on petabytes of data in Hadoop clusters.
Rust bindings and safe wrapper APIs for Hadoop's libhdfs, enabling HDFS access from Rust applications.
A Scala/Spark library for efficient processing, extraction, and derivation of web archive data (CDX/WARC).
A lightweight, HDFS-compatible file system built over Cassandra with a fat driver design for easy deployment.
Provides HBase adapters for reading and writing data within Cascading data processing workflows on Hadoop clusters.
A splitable Hadoop InputFormat for processing concatenated GZIP files and web archive (*.warc.gz) data efficiently in distributed systems.
A Hadoop/MapReduce tool that splits and partitions web archive records in (W)ARC files by MIME type and year.
Open-Awesome is built by the community, for the community. Submit a project, suggest an awesome list, or help improve the catalog on GitHub.