Showing 31 of 103 projects
A free, open-source alternative to Spark UI and Spark History Server with enhanced CPU and memory metrics visualizations.
A curated list of the best resources, tools, libraries, and documentation for the Apache Cassandra database ecosystem.
A curated collection of resources and guides for understanding, selecting, and using NoSQL databases effectively.
A Python SQL parser that converts SQL queries into JSON-izable parse trees for translation to non-SQL datastores.
An idiomatic Clojure dataframe library that runs on Apache Spark, providing a seamless interface for data processing and machine learning.
An open-source unit test framework for Hive SQL queries, enabling TDD without installed dependencies via JUnit 4 and 5.
A Go library and CLI tool for validating CSV files against RFC 4180 standards.
A DataOps-friendly data quality monitoring platform with customizable checks, dashboards, and incident management for multiple data sources.
A curated list of awesome HBase projects, clients, frameworks, tools, and resources.
A visual development platform for building, deploying, and managing streaming analytics applications with multiple engine bindings.
A Spark library for reading from and writing to Google BigQuery using DataFrames and SQL.
A Rust DataFrame and data engineering library with PySpark/SQL-like syntax, built for business data pipelines with Microsoft stack integration.
A manifesto advocating for treating database interactions, queries, and lifecycle management as plain code with SQL as the primary language.
An experimental Rust client for Apache Spark Connect, providing a DataFrame API to interact with Spark clusters.
A Python framework for building and deploying serverless data and ML pipelines on AWS using AWS CDK.
A PHP client extension for the TDengine big data engine, with Swoole coroutine support.
A simple utility for testing Apache Hive scripts locally without requiring Java development skills.
An easy-to-use Python feature store for machine learning, optimized for timeseries data and built on Dask.
A pure PHP library for reading and writing Parquet columnar storage files without external dependencies.
A Julia client interface for reading from and writing to the TypeDB knowledge graph database.
A native Common Lisp interface for the InfluxDB time series database.
A native Common Lisp interface for the InfluxDB time series database.
A vendor-neutral, declarative data quality engine that defines checks in YAML and runs anywhere.
A lightweight, HDFS-compatible file system built over Cassandra with a fat driver design for easy deployment.
A weekly online meetup and resource hub for Apache Cassandra topics, featuring talks, tutorials, and community discussions.
A pure Go toolkit for data engineering and classic machine learning with zero external dependencies.
Extract, transform, and load (ETL) scripts for exporting and streaming EOS blockchain data.
A blazingly fast data comparison tool for Python that instantly compares massive CSV/Parquet datasets, powered by Rust.
A Python package for converting PCAP network capture files to Parquet, CSV, or JSON formats.
A CRDT-based merge library that guarantees mathematical convergence for DataFrames, JSON, ML models, and distributed agents.
A Zsh plugin that enhances Databricks CLI with convenient aliases, profile management, and job run analysis using the 'dbrs' prefix.
Open-Awesome is built by the community, for the community. Submit a project, suggest an awesome list, or help improve the catalog on GitHub.