Open-Awesome
CategoriesAlternativesStacksSelf-HostedExplore
Open-Awesome

© 2026 Open-Awesome. Curated for the developer elite.

TermsPrivacyAboutGitHubRSS
  1. Home
  2. Machine Learning
  3. Aqueduct

Aqueduct

Apache-2.0Gov0.3.6

An open-source MLOps framework for defining and deploying machine learning and LLM workloads across any cloud infrastructure.

Visit WebsiteGitHubGitHub
517 stars20 forks0 contributors

What is Aqueduct?

Aqueduct is an open-source MLOps framework that enables data scientists and engineers to define machine learning and LLM workflows in vanilla Python and deploy them on any cloud infrastructure, such as Kubernetes, Spark, or AWS Lambda. It addresses the fragmented nature of MLOps by providing a unified interface to existing tools while ensuring centralized visibility into pipeline execution. The framework allows seamless movement of code between different cloud layers without requiring a rip-and-replace approach.

Target Audience

Data scientists and ML engineers who need to deploy and manage machine learning or LLM workflows across diverse cloud infrastructure like Kubernetes, Spark, Airflow, or AWS Lambda. It is suited for teams dealing with the complexity of siloed MLOps tools and seeking a Python-native solution.

Value Proposition

Developers choose Aqueduct because it offers a Python-native API without DSLs or YAML configurations, enabling quick production deployment. Its unique selling point is the ability to integrate with and orchestrate workflows across multiple existing cloud infrastructure systems while providing centralized visibility into code, data, and metadata for reliability and debugging.

Overview

Aqueduct is no longer being maintained. Aqueduct allows you to run LLM and ML workloads on any cloud infrastructure.

Use Cases

Best For

  • Deploying machine learning or LLM workflows across heterogeneous cloud environments (e.g., combining Kubernetes and AWS Lambda in a single pipeline).
  • Teams seeking a unified MLOps framework to manage siloed infrastructure tools without replacing their existing investments.
  • Data scientists who want to write workflows in vanilla Python and avoid learning new DSLs or complex YAML configurations.
  • Organizations requiring centralized visibility and monitoring of ML pipeline execution, including code, data, metrics, and metadata.
  • Running secure ML workloads entirely within a private cloud environment to ensure data and code security.
  • Orchestrating and scheduling ML workflows on-demand or automatically across distributed systems like Spark or Airflow.

Not Ideal For

  • Projects requiring a fully managed ML platform with built-in data versioning and experiment tracking out-of-the-box.
  • Teams already deeply integrated with a single, monolithic MLOps ecosystem like Kubeflow that meets all their needs.
  • Simple deployments running entirely on one infrastructure type (e.g., only Kubernetes) where lightweight tools suffice.

Pros & Cons

Pros

Python-Native API

Enables defining ML and LLM workflows in vanilla Python without DSLs or YAML, as shown in the quickstart guide for rapid deployment and ease of use.

Multi-Cloud Orchestration

Integrates seamlessly with diverse infrastructures like Kubernetes, Spark, and AWS Lambda, allowing tasks to run across different systems without replacing existing tooling.

Centralized Visibility

Provides a unified UI to monitor code, data, metrics, and metadata from each workflow run, enhancing reliability and debugging capabilities as illustrated in the README screenshot.

Secure Cloud Execution

Runs entirely within the user's cloud environment, ensuring data and code security without relying on external services, aligning with the emphasis on security in the features list.

Cons

No Built-In Data Management

Lacks native features for data versioning and experiment tracking, requiring users to integrate and manage additional tools for a complete MLOps stack, which adds complexity.

Complex Multi-Cloud Setup

Configuring and managing connections to various cloud infrastructures like Kubernetes and AWS Lambda can be challenging, especially for teams without extensive DevOps expertise.

Limited Community and Ecosystem

As a newer framework, it has fewer pre-built operators, integrations, and community support compared to established platforms like Airflow or MLflow, potentially slowing adoption.

Frequently Asked Questions

Quick Stats

Stars517
Forks20
Contributors0
Open Issues10
Last commit3 years ago
CreatedSince 2022

Tags

#cloud-infrastructure#ai#open-source#workflow-orchestration#data-science#kubernetes#llm#python3#mlops#python#ml#data-pipelines#data#machine-learning

Built With

K
Kubernetes
S
SPARK
P
Python
A
AWS Lambda

Links & Resources

Website

Included in

Machine Learning72.2k
Auto-fetched 5 hours ago

Related Projects

promptfoopromptfoo

Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.

Stars23,550
Forks2,117
Last commit7 hours ago
dvcdvc

🦉 Data Versioning and ML Experiments

Stars15,771
Forks1,314
Last commit3 days ago
txtaitxtai

💡 All-in-one AI framework for semantic search, LLM orchestration and language model workflows

Stars12,749
Forks850
Last commit18 hours ago
KedroKedro

Kedro is a toolbox for production-ready data science. It uses software engineering best practices to help you create data engineering and data science pipelines that are reproducible, maintainable, and modular.

Stars10,931
Forks1,058
Last commit18 hours ago
Community-curated · Updated weekly · 100% open source

Found a gem we're missing?

Open-Awesome is built by the community, for the community. Submit a project, suggest an awesome list, or help improve the catalog on GitHub.

Submit a projectStar on GitHub