Open-Awesome
CategoriesAlternativesStacksSelf-HostedExplore
Open-Awesome

© 2026 Open-Awesome. Curated for the developer elite.

TermsPrivacyAboutGitHubRSS
  1. Home
  2. Data Engineering
  3. Gobblin

Gobblin

Apache-2.0Javagobblin_0.11.0

A distributed data integration framework for big data ecosystems, handling ingestion, replication, organization, and lifecycle management for both streaming and batch data.

Visit WebsiteGitHubGitHub
2.3k stars749 forks0 contributors

What is Gobblin?

Apache Gobblin is a distributed data integration framework designed to simplify big data integration tasks such as data ingestion, replication, organization, and lifecycle management. It handles both streaming and batch data ecosystems, providing a scalable solution for managing structured and byte-oriented data across heterogeneous environments. The framework is optimized for ELT patterns with inline transformations and is battle-tested at petabyte scale in production environments.

Target Audience

Data engineers and architects working with large-scale data ecosystems who need reliable ingestion, replication, and lifecycle management across diverse data sources and sinks. Organizations with complex data integration requirements across Hadoop, cloud storage, and external APIs will benefit most.

Value Proposition

Developers choose Apache Gobblin for its proven scalability in production environments, comprehensive feature set for data management, and flexibility in supporting both stream and batch execution modes. Its ability to handle petabyte-scale workflows while providing fault tolerance, data quality checking, and compliance management makes it a robust alternative to building custom integration solutions.

Overview

A distributed data integration framework that simplifies common aspects of big data integration such as data ingestion, replication, organization and lifecycle management for both streaming and batch data ecosystems.

Use Cases

Best For

  • Stream and batch ingestion from Kafka to data lakes like HDFS, S3, or ADLS
  • Bulk-loading serving stores from data lakes (e.g., HDFS to Couchbase)
  • Data synchronization across federated data lakes (HDFS to S3, S3 to ADLS)
  • Integrating external vendor APIs (Salesforce, Dynamics) with data stores
  • Enforcing data retention policies and GDPR compliance deletions
  • Managing data organization tasks like compaction, partitioning, and deduplication

Not Ideal For

  • Projects requiring complex data transformations or heavy ETL processing beyond inline ELT patterns
  • Teams already standardized on a comprehensive workflow orchestrator like Airflow for all scheduling and dependency management
  • Organizations with small-scale or homogeneous data environments where lighter tools or custom scripts suffice
  • Real-time applications needing sub-second latency, as Gobblin's streaming is optimized for reliability over ultra-low latency

Pros & Cons

Pros

Production-Proven Scalability

Battle-tested at petabyte-scale by companies like LinkedIn and PayPal, ensuring reliability for large data workflows as highlighted in the README.

Comprehensive Data Management

Offers end-to-end capabilities including ingestion, compaction, deduplication, and lifecycle management, covering complex data integration needs from the README.

Flexible Execution Modes

Supports both stream and batch execution, allowing adaptable data processing workflows, as noted in the highlights section.

Robust Fault Tolerance

Includes features like task partitioning, state management, and atomic data publishing, enhancing reliability in distributed environments per the README.

Cons

Limited Transformation Engine

Delegates complex data processing to external systems like Spark, adding dependency and overhead, as admitted in the 'Apache Gobblin is NOT' section.

Complex Initial Setup

Building from source requires Gradle and Maven with non-trivial instructions, which can be daunting for new users, as seen in the build requirements.

Dependency on Mature Ecosystems

Best suited for Hadoop or cloud storage environments, limiting appeal for modern, cloud-native setups without extensive integration work.

Incubator Project Risks

Being in the Apache Incubator may imply ongoing development and potential breaking changes, affecting stability for production use.

Frequently Asked Questions

Quick Stats

Stars2,270
Forks749
Contributors0
Open Issues0
Last commit29 days ago
CreatedSince 2014

Tags

#stream-processing#apache#batch-processing#replication#data-integration#distributed-systems#data-replication#data-lake#management#big-data#data#data-ingestion#etl-framework

Built With

J
Java
D
Docker
G
Gradle

Links & Resources

Website

Included in

Data Engineering8.5k
Auto-fetched 4 hours ago

Related Projects

kafka-managerkafka-manager

CMAK is a tool for managing Apache Kafka clusters

Stars11,926
Forks2,476
Last commit3 years ago
KreuzbergKreuzberg

A polyglot document intelligence framework with a Rust core. Extract text, metadata, images, and structured information from PDFs, Office documents, images, and 97+ formats. Available for Rust, Python, Ruby, Java, Go, PHP, Elixir, C#, R, C, TypeScript (Node/Bun/Wasm/Deno)- or use via CLI, REST API, or MCP server.

Stars8,691
Forks526
Last commit4 hours ago
kafka-dockerkafka-docker

Dockerfile for Apache Kafka

Stars6,963
Forks2,662
Last commit2 years ago
kafkacatkafkacat

Generic command line non-JVM Apache Kafka producer and consumer

Stars5,768
Forks500
Last commit2 years ago
Community-curated · Updated weekly · 100% open source

Found a gem we're missing?

Open-Awesome is built by the community, for the community. Submit a project, suggest an awesome list, or help improve the catalog on GitHub.

Submit a projectStar on GitHub