Open-Awesome
CategoriesAlternativesStacksSelf-HostedExplore
Open-Awesome

© 2026 Open-Awesome. Curated for the developer elite.

TermsPrivacyAboutGitHubRSS
  1. Home
  2. Tags
  3. Etl

Etl

106 projects

Showing 36 of 106 projects

Airflow
AirflowPython

A platform to programmatically author, schedule, and monitor workflows as code.

#apache#airflow#devops
Stars46.2k
Forks17.4k
Last commit18 hours ago
Apache AirFlow
Apache AirFlowPython

A platform to programmatically author, schedule, and monitor workflows as code.

#apache#airflow#devops
Stars46.2k
Forks17.4k
Last commit18 hours ago
Airbyte (k)
Airbyte (k)Python

Open-source data integration platform for building ELT pipelines from APIs, databases, and files to data warehouses, lakes, and lakehouses.

#open-source#pipeline#data-integration
Stars21.7k
Forks5.3k
Last commit19 hours ago
Dagster
DagsterPython

An orchestration platform for developing, deploying, and monitoring data pipelines and assets.

#data-orchestration#data-assets#devops
Stars15.9k
Forks2.2k
Last commit1 day ago
Logstash
LogstashJava

A server-side data processing pipeline that ingests, transforms, and ships logs and events from multiple sources.

#event-processing#jruby#server-side
Stars14.9k
Forks3.5k
Last commit1 day ago
dbt-core
dbt-coreRust

A transformation tool that enables data analysts and engineers to transform data using software engineering best practices.

#version-control#pypa#business-intelligence
Stars13.5k
Forks2.5k
Last commit19 hours ago
Debezium (k)
Debezium (k)Java

A low-latency platform for change data capture (CDC) that streams row-level changes from databases to applications.

#database#event-driven-architecture#cqrs
Stars12.9k
Forks3.0k
Last commit1 day ago
CocoIndex
CocoIndexRust

An ultra-performant data transformation framework for AI, with incremental processing and data lineage built-in.

#data-indexing#semantic-search#data-lineage
Stars11.0k
Forks848
Last commit1 day ago
awesome-data-engineering
awesome-data-engineering

A curated list of data engineering tools, frameworks, databases, and resources for software developers.

#stream-processing#workflow-orchestration#awesome-list
Stars8.9k
Forks1.6k
Last commit6 days ago
Benthos
BenthosGo

A high-performance, declarative stream processor that connects various sources and sinks with built-in data transformation capabilities.

#stream-processing#cqrs#message-queue
Stars8.7k
Forks952
Last commit23 hours ago
Benthos
BenthosGo

A high-performance, resilient stream processor that connects various sources and sinks, performs data transformations, and guarantees at-least-once delivery.

#stream-processing#declarative-config#cqrs
Stars8.7k
Forks952
Last commit23 hours ago
Pentaho Data Integration (.3k)
Pentaho Data Integration (.3k)Java

An open-source ETL (Extract, Transform, Load) tool for data integration and migration.

#plugin-system#data-integration#business-intelligence
Stars8.4k
Forks3.6k
Last commit19 hours ago
pgloader
pgloaderCommon Lisp

A data loading and migration tool for PostgreSQL that handles errors gracefully and transforms data from various sources.

#mssql#migration#database
Stars6.5k
Forks607
Last commit9 days ago
CloudQuery
CloudQueryGo

Open-source data pipelines to sync cloud infrastructure metadata from AWS, Azure, GCP, and 70+ sources into your data warehouse.

#sql-queryable#multi-cloud#apache-arrow
Stars6.5k
Forks548
Last commit6 days ago
CloudQuery
CloudQueryGo

Open-source data pipelines for cloud asset inventory, CSPM, FinOps, and vulnerability management across AWS, Azure, GCP, and 70+ sources.

#sql-queryable#multi-cloud#apache-arrow
Stars6.5k
Forks548
Last commit6 days ago
Apache NiFi (k)
Apache NiFi (k)Java

An easy-to-use, powerful, and reliable system to process and distribute data across cybersecurity, observability, and AI pipelines.

#hacktoberfest#apache#observability
Stars6.2k
Forks3.0k
Last commit1 day ago
Azkaban (.5k)
Azkaban (.5k)Java

Azkaban is a batch workflow job scheduler created at LinkedIn to manage Hadoop jobs.

#hacktoberfest#gradle#batch-processing
Stars4.5k
Forks1.6k
Last commit2 years ago
Dedupe
DedupePython

A Python library using machine learning for accurate and scalable fuzzy matching, record deduplication, and entity resolution on structured data.

#data-cleaning#de duplicating#python-library
Stars4.5k
Forks575
Last commit1 year ago
Maxwell's daemon (.2k)
Maxwell's daemon (.2k)Java

A MySQL change data capture daemon that streams database changes as JSON to Kafka, Kinesis, and other platforms.

#change-data-capture#database-replication#kafka
Stars4.3k
Forks1.0k
Last commit6 days ago
aws-sdk-pandas
aws-sdk-pandasPython

A Python library that simplifies data integration between pandas and AWS services like Athena, S3, Redshift, and more.

#apache-arrow#data-science#glue-catalog
Stars4.1k
Forks737
Last commit3 days ago
aws-data-wrangler
aws-data-wranglerPython

A Python library that simplifies data integration between pandas and AWS services like Athena, S3, Redshift, and more.

#apache-arrow#data-science#redshift
Stars4.1k
Forks737
Last commit3 days ago
ingestr
ingestrGo

A CLI tool to copy data between any databases and platforms with a single command, no code required.

#dlt#mssql#no-code
Stars3.8k
Forks143
Last commit1 day ago
QSV
QSVRust

A blazing-fast command-line toolkit for querying, slicing, analyzing, transforming, and validating tabular data (CSV, Excel, JSONL, etc.).

#ckan#parquet#luau
Stars3.7k
Forks104
Last commit19 hours ago
awesome-etl list
awesome-etl list

A curated list of awesome ETL frameworks, libraries, and software for data integration and pipeline development.

#open-source#workflow-orchestration#data-integration
Stars3.6k
Forks372
Last commit2 months ago
pgsync
pgsyncRuby

A command-line tool to efficiently and securely sync data between PostgreSQL databases with parallel transfers and data masking.

#devops#backup-tool#ruby-gem
Stars3.5k
Forks218
Last commit7 months ago
Koalas
KoalasPython

Koalas provides the pandas DataFrame API on Apache Spark, enabling data scientists to work with big data using familiar pandas syntax.

#apache-spark#spark#mlflow
Stars3.4k
Forks369
Last commit2 years ago
elasticsearch-jdbc
elasticsearch-jdbcJava

A Java-based tool for importing tabular data from JDBC sources into Elasticsearch for indexing.

#database#bulk-indexing#search-index
Stars2.8k
Forks698
Last commit1 month ago
amazon-redshift-utils
amazon-redshift-utilsPython

A collection of utilities, scripts, and views for managing, optimizing, and automating Amazon Redshift data warehouse operations.

#sql-scripts#performance-tuning#data-migration
Stars2.8k
Forks1.2k
Last commit
Scio
ScioScala

A Scala API for Apache Beam and Google Cloud Dataflow, enabling unified batch and streaming data processing.

#stream-processing#batch-processing#batch
Stars2.6k
Forks532
Last commit2 days ago
Hamilton
HamiltonJupyter Notebook

A Python library for defining portable, modular, and testable data transformation DAGs with built-in lineage and metadata.

#data-lineage#etl-pipeline#python-library
Stars2.6k
Forks201
Last commit5 days ago
Hamilton
HamiltonJupyter Notebook

A lightweight Python library for creating portable, expressive, and testable data transformation DAGs with built-in lineage and metadata.

#data-lineage#etl-pipeline#python-library
Stars2.6k
Forks201
Last commit5 days ago
VDP
VDPPython

Instill Core is a full-stack AI infrastructure tool for data, model, and pipeline orchestration to build versatile AI-first applications.

#hacktoberfest#ai#ai-infrastructure
Stars2.3k
Forks124
Last commit1 month ago
Proton
ProtonC++

A single C++ binary SQL engine for high-performance stream processing, analytics, observability, and AI/ML pipelines.

#stream-processing#sql-engine#iceberg
Stars2.2k
Forks110
Last commit21 days ago
go-streams
go-streamsGo

A lightweight and efficient stream processing library for Go, providing a declarative DSL to build data pipelines.

#stream-processing#pulsar#redis
Stars2.2k
Forks174
Last commit6 months ago
sqlite-utils
sqlite-utilsPython

A Python CLI utility and library for manipulating SQLite databases, including importing JSON/CSV and running in-memory queries.

#click#sqlite-database#datasette-tool
Stars2.1k
Forks157
Last commit11 days ago
Onyx
OnyxClojure

A masterless, cloud-scale, fault-tolerant distributed computation system for batch and stream processing written in Clojure.

#stream-processing#batch-processing#distributed
Stars2.1k
Forks200
Last commit6 years ago
Page 1 of 3Next

Related Tags

Community-curated · Updated weekly · 100% open source

Found a gem we're missing?

Open-Awesome is built by the community, for the community. Submit a project, suggest an awesome list, or help improve the catalog on GitHub.

Submit a projectStar on GitHub
10 months ago
#Data Integration26
#Data Pipeline26
#Data Processing23
#Data Engineering23
#Big Data22
#Python19
#Java18
#Data Pipelines15
#Batch Processing15
#Apache Spark14
#Stream Processing14
#Data Transformation13