Open-Awesome
CategoriesAlternativesStacksSelf-HostedExplore
Open-Awesome

© 2026 Open-Awesome. Curated for the developer elite.

TermsPrivacyAboutGitHubRSS
  1. Home
  2. Tags
  3. Data Pipeline

Data Pipeline

131 projects

Showing 36 of 131 projects

Apache Kafka Streams
Apache Kafka StreamsJava

A distributed event streaming platform for building high-performance data pipelines, streaming analytics, and data integration.

#stream-processing#message-queue#data-integration
Stars33.7k
Forks15.5k
Last commit22 hours ago
Vector
VectorRust

A high-performance, end-to-end observability data pipeline for collecting, transforming, and routing logs and metrics.

#stream-processing#hacktoberfest#pipelines
Stars22.5k
Forks2.3k
Last commit3 days ago
vector
vectorRust

A high-performance, end-to-end observability data pipeline for collecting, transforming, and routing logs and metrics.

#stream-processing#hacktoberfest#pipelines
Stars22.5k
Forks2.3k
Last commit3 days ago
privacy-preserving ML
privacy-preserving ML

A curated list of awesome open-source libraries for deploying, monitoring, versioning, and scaling production machine learning systems.

#explainability#deep-learning#interpretability
Stars20.9k
Forks2.6k
Last commit1 day ago
Awesome Production Machine Learning
Awesome Production Machine Learning

A curated list of awesome open-source libraries for deploying, monitoring, versioning, and scaling production machine learning systems.

#ai-infrastructure#open-source#explainability
Stars20.9k
Forks2.6k
Last commit1 day ago
Telegraf PostgreSQL plugin
Telegraf PostgreSQL pluginGo

A plugin-driven agent for collecting, processing, aggregating, and writing metrics, logs, and arbitrary data.

#plugin-system#observability#logs
Stars17.8k
Forks5.8k
Last commit13 hours ago
Logstash
LogstashJava

A server-side data processing pipeline that ingests, transforms, and ships logs and events from multiple sources.

#event-processing#jruby#server-side
Stars14.9k
Forks3.5k
Last commit23 hours ago
awesome-bigdata
awesome-bigdata

A curated list of awesome big data frameworks, resources, and tools across various categories.

#database#data-science#distributed-systems
Stars14.6k
Forks2.6k
Last commit1 month ago
Fluentd
FluentdRuby

An open-source log collector that unifies logging infrastructure by collecting events from various sources and routing them to multiple destinations.

#event-processing#devops#observability
Stars13.6k
Forks1.4k
Last commit1 day ago
Debezium (k)
Debezium (k)Java

A low-latency platform for change data capture (CDC) that streams row-level changes from databases to applications.

#database#event-driven-architecture#cqrs
Stars13.1k
Forks3.0k
Last commit18 hours ago
Machine Learning Interviews
Machine Learning InterviewsHTML

A practical booklet covering the four main steps of designing machine learning systems with 27 interview questions.

#data-science#machine-learning-production#production-ml
Stars10.6k
Forks1.6k
Last commit
Kreuzberg
KreuzbergRust

A polyglot document intelligence framework with a Rust core for extracting text, metadata, and structured data from 91+ file formats.

#text-extraction#document-intelligence#batch-processing
Stars9.3k
Forks582
Last commit10 hours ago
Benthos
BenthosGo

A high-performance, resilient stream processor that connects various sources and sinks, performs data transformations, and guarantees at-least-once delivery.

#stream-processing#declarative-config#cqrs
Stars8.7k
Forks964
Last commit10 hours ago
Benthos
BenthosGo

A high-performance, declarative stream processor that connects various sources and sinks with built-in data transformation capabilities.

#stream-processing#cqrs#message-queue
Stars8.7k
Forks964
Last commit10 hours ago
Pentaho Data Integration (.3k)
Pentaho Data Integration (.3k)Java

An open-source ETL (Extract, Transform, Load) tool for data integration and migration.

#plugin-system#data-integration#business-intelligence
Stars8.4k
Forks3.6k
Last commit9 hours ago
snowplow
snowplowScala

Open-source customer data infrastructure that collects, validates, and enriches behavioral event data for AI and analytics.

#snowplow-events#data-warehouse-integration#event-tracking
Stars7.0k
Forks1.2k
Last commit2 months ago
CloudQuery
CloudQueryGo

Open-source data pipelines to sync cloud infrastructure metadata from AWS, Azure, GCP, and 70+ sources into your data warehouse.

#sql-queryable#multi-cloud#apache-arrow
Stars6.5k
Forks557
Last commit1 day ago
CloudQuery
CloudQueryGo

Open-source data pipelines for cloud asset inventory, CSPM, FinOps, and vulnerability management across AWS, Azure, GCP, and 70+ sources.

#sql-queryable#multi-cloud#apache-arrow
Stars6.5k
Forks557
Last commit1 day ago
Apache NiFi (k)
Apache NiFi (k)Java

An easy-to-use, powerful, and reliable system to process and distribute data across cybersecurity, observability, and AI pipelines.

#hacktoberfest#apache#observability
Stars6.2k
Forks3.0k
Last commit11 hours ago
kcat (.7k)
kcat (.7k)C

A lightweight command-line tool for producing, consuming, and inspecting Apache Kafka messages, similar to netcat for Kafka.

#devops#message-queue#command-line-tool
Stars5.8k
Forks501
Last commit2 years ago
kafkacat
kafkacatC

A lightweight, non-JVM command-line tool for producing, consuming, and inspecting Apache Kafka messages.

#devops#message-queue#command-line-tool
Stars5.8k
Forks501
Last commit2 years ago
fluvio
fluvioRust

A distributed data streaming engine with stateful stream processing for building responsive data-intensive applications.

#stream-processing#event-driven#webassembly
Stars5.2k
Forks530
Last commit8 days ago
Clidey WhoDB
Clidey WhoDBGo

A lightweight, fast, and beautiful database management tool with AI-powered chat interface for PostgreSQL, MySQL, SQLite, MongoDB, Redis, and more.

#data-lineage#database#explorer
Stars5.0k
Forks238
Last commit17 hours ago
RudderStack
RudderStackGo

An open-source, privacy-focused customer data platform (CDP) that collects, processes, and routes event data to warehouses and tools.

#event-collection#segment-alternative#warehouse-management
Stars4.5k
Forks64
Last commit9 hours ago
DotnetSpider
DotnetSpiderC#

A lightweight, efficient, and fast high-level web crawling and scraping framework for .NET.

#web-crawling#distributed#redis
Stars4.1k
Forks1.1k
Last commit5 months ago
ingestr
ingestrGo

A CLI tool to copy data between any databases and platforms with a single command, no code required.

#dlt#mssql#no-code
Stars4.0k
Forks152
Last commit14 hours ago
Dagu
DaguGo

Self-hostable workflow orchestrator for teams whose main work isn't orchestration. Declarative YAML over your scripts, SSH commands, containers, etc; keep workflows separate from business logic. One binary, no database, runs on limited H/W resources. Alternative to Airflow / Cron / Job Scheduler.

#task-automation#devops#task-scheduler
Stars3.8k
Forks320
Last commit
Dagu
DaguGo

A local-first, single-binary workflow orchestration engine that runs declarative DAGs from laptop to distributed cluster.

#task-automation#devops#task-scheduler
Stars3.8k
Forks320
Last commit2 days ago
tengo
tengoGo

A fast, embeddable scripting language for Go applications, compiled to bytecode and executed on a stack-based VM.

#programming-language#compiler#rules-engine
Stars3.8k
Forks337
Last commit4 months ago
databus
databusJava

A source-agnostic distributed change data capture system for reliably capturing and streaming primary data changes.

#linkedin#oracle#change-data-capture
Stars3.7k
Forks736
Last commit2 years ago
Ensemble-Strategy
Ensemble-StrategyPython

An AI-native modular infrastructure for quantitative trading, featuring a weight-centric architecture for building, testing, and deploying algorithmic strategies.

#backtesting#algorithmic-trading#finrl
Stars3.7k
Forks1.1k
Last commit
Streaming
StreamingJavaScript

A curated list of awesome streaming frameworks, applications, readings, and resources for stream processing.

#stream-processing#message-queue#real-time-analytics
Stars3.0k
Forks324
Last commit14 hours ago
Numaflow
NumaflowRust

A Kubernetes-native, serverless platform for running massively parallel data and streaming jobs with exactly-once semantics.

#stream-processing#hacktoberfest#event-driven-architecture
Stars2.8k
Forks172
Last commit19 hours ago
broadway
broadwayElixir

Build concurrent, multi-stage data ingestion and processing pipelines with Elixir, supporting back-pressure, batching, and fault tolerance.

#event-driven#back-pressure#elixir
Stars2.7k
Forks179
Last commit1 month ago
Scio
ScioScala

A Scala API for Apache Beam and Google Cloud Dataflow, enabling unified batch and streaming data processing.

#stream-processing#batch-processing#batch
Stars2.6k
Forks533
Last commit12 days ago
Proton
ProtonC++

A single C++ binary SQL engine for high-performance stream processing, analytics, observability, and AI/ML pipelines.

#stream-processing#sql-engine#iceberg
Stars2.3k
Forks109
Last commit6 days ago
Page 1 of 4

Related Tags

Community-curated · Updated weekly · 100% open source

Found a gem we're missing?

Open-Awesome is built by the community, for the community. Submit a project, suggest an awesome list, or help improve the catalog on GitHub.

Submit a projectStar on GitHub
3 years ago
2 days ago
4 months ago
Next
#Stream Processing30
#Etl26
#Data Integration18
#Kafka18
#Big Data17
#Docker16
#Data Processing14
#Monitoring14
#Distributed Systems13
#Python13
#Java12
#Data Ingestion12