Showing 23 of 131 projects
A tool for collecting, validating, enhancing, and merging GTFS feeds for public transport data, with a focus on German agencies.
A lightweight ETL library and data integration toolbox for .NET, enabling programmatic data flow pipelines.
A public dataset of Ethereum network events including beacon chain, mempool, and canonical chain data for analysis.
Combines realtime transit information from multiple sources into single output files for public transportation systems.
A web SDK for collecting and reporting user clickstream events from browsers to AWS analytics pipelines.
A Scala sample application demonstrating Spark Streaming integration with Kafka and Cassandra for data processing.
A distributed data stream pipeline for querying, augmenting, and transforming data using Elixir pattern-matching rules.
A Suricata plugin that outputs Eve JSON events to Apache Kafka for real-time network security monitoring.
A service that consumes, transforms, and republishes JSON messages on MQTT topics using configurable JSON-e templates.
A Python program that fetches seismic waveform data from IRIS and writes it to a TDengine time-series database.
A sample application demonstrating how to consume Amazon Cognito streams and model data in Amazon Redshift.
A stream processing platform with a pluggable architecture for connecting data feeds to transformation operators.
Extract, transform, and load (ETL) scripts for exporting and streaming EOS blockchain data.
A Python data pipeline for processing GTFS static and realtime bus datasets from Transport for NSW.
An Apache Flume source plugin for ingesting UDP messages directly into Flume data pipelines.
A JSON-based language and library for transforming JSON data using declarative operations.
Change data capture from PostgreSQL into Kafka using logical decoding, enabling real-time data streaming.
AWS_MSK_IAM authentication plugin for Broadway Kafka, enabling Elixir applications to connect to Amazon MSK using IAM credentials.
A Node.js tool that subscribes to MQTT topics and forwards messages to Elasticsearch for indexing and analysis.
A Kafka consumer that archives data streams into TDEngine time-series database for efficient storage and querying.
An Embulk output plugin for writing data to InfluxDB time-series databases.
A lightweight Node.js ETL framework for extracting data from databases and loading it into data lakes and warehouses.
A dbt/Evidence project for tracking and analyzing personal coffee ratings with automated data pipelines.
Open-Awesome is built by the community, for the community. Submit a project, suggest an awesome list, or help improve the catalog on GitHub.