Showing 34 of 106 projects
A distributed, scalable database built for stream processing applications on Apache Kafka using SQL syntax.
A Spark library for reading and writing data between Spark SQL and MongoDB collections.
An idiomatic Clojure dataframe library that runs on Apache Spark, providing a seamless interface for data processing and machine learning.
A Go library for declarative JSON-to-JSON transformations using JSON specifications.
A collection of import/export commands for the Neo4j shell to load and dump graph data in various formats.
A collection of connectors enabling Apache HBase integration with Kafka, Spark, and other data processing systems.
A mapping language and engine for converting complex, nested data between schemas, with extensibility via plugins.
R client for the Elasticsearch HTTP API, enabling data indexing, search, and analysis from R.
A Go-based toolkit for fast ETL and feature extraction on Hadoop, optimized for rapid development and execution.
A Go-based toolset for data extraction, transformation, and loading, providing powerful data synchronization capabilities.
A Spark library for reading from and writing to Google BigQuery using DataFrames and SQL.
Official Neo4j JDBC Driver
Operator and codec library for building real-time streaming applications on Apache Apex.
A Java library for enriching, transforming, and filtering JSON documents using configurable pipelines.
An experimental Rust client for Apache Spark Connect, providing a DataFrame API to interact with Spark clusters.
A MongoDB to Neo4j document manager for live one-way synchronization, enabling polyglot persistence by converting documents into a graph structure.
A Spark application for migrating data to ScyllaDB from CQL-compatible databases or DynamoDB via Alternator.
A pure PHP library for reading and writing Parquet columnar storage files without external dependencies.
A command-line tool to clean and normalize CSV files by removing invalid records and standardizing data.
A Go library for Apache Avro with strong typing, SQL integration, and Redshift schema generation.
A modular framework for ingesting and processing Algorand blockchain data into external applications.
A lightweight ETL library and data integration toolbox for .NET, enabling programmatic data flow pipelines.
A dynamic framework for processing high-volume data streams with subsecond pipeline instantiation and modification latency.
A Java-based high-performance importer skeleton for complex, business-logic-heavy data imports into Neo4j from SQL databases, CSV files, and other sources.
A PHP library for live importing Google Sheets data into data warehouses with periodic delta loads.
A configurable GenStage-based bulk processor for efficiently inserting data into Elasticsearch from Elixir applications.
A pure Go toolkit for data engineering and classic machine learning with zero external dependencies.
Extract, transform, and load (ETL) scripts for exporting and streaming EOS blockchain data.
A fast, memory-efficient CLI tool for splitting large SQL dump files into individual table files and converting between SQL dialects.
A blazingly fast data comparison tool for Python that instantly compares massive CSV/Parquet datasets, powered by Rust.
An Embulk output plugin for writing data to InfluxDB time-series databases.
Documentation for Sitecore Data Exchange Framework, an ETL tool for Sitecore.
A Java project that reads from and writes to TDengine using Apache Spark for data processing.
Import CSV files from AWS S3 into Cassandra using Apache Spark with a simple configuration-based approach.
Open-Awesome is built by the community, for the community. Submit a project, suggest an awesome list, or help improve the catalog on GitHub.