Showing 36 of 624 projects
A curated collection of open data sources across government, academic, and private sectors for data science and research.
An open-source MLOps framework for defining and deploying machine learning and LLM workloads across any cloud infrastructure.
A lightweight Python tool for generating rich summary statistics of pandas and Polars dataframes directly in the console.
An AutoML framework that generates and customizes machine learning pipelines using declarative JSON-AI syntax.
A Julia package providing metaprogramming macros to simplify DataFrame manipulation with a more concise syntax.
A high-performance data profiler for discovering and validating complex patterns in datasets, enabling data cleaning and quality analysis.
A high-performance data profiler for discovering and validating complex patterns like functional dependencies, inclusion dependencies, and association rules.
A pure Go library for making predictions with Gradient Boosting Regression Trees models from LightGBM, XGBoost, and scikit-learn.
An overlay companion for pandas that provides real-time hints and tips to improve data analysis code.
An open-source machine learning solution for the Home Credit Default Risk Kaggle competition, providing reproducible code and experiments.
F# kernel for Jupyter notebooks, enabling interactive data science and exploration with F#.
A SQL GUI extension for JupyterLab that enables point-and-click database exploration and query execution.
A Python library for class-imbalanced ensemble learning with 30+ algorithms, built on scikit-learn.
Jupyter notebooks implementing algorithms, proofs, and summaries from 'The Elements of Statistical Learning' textbook.
A Python meta-library for community detection in complex networks, implementing algorithms, fitness functions, and visualization.
A scikit-learn-compatible Python implementation of ReBATE, a suite of Relief-based feature selection algorithms for machine learning.
A JupyterLab extension that adds support for creating notebooks from customizable templates.
A practical demo using LSTM neural networks with TensorFlow to predict lottery numbers.
A curated list of research, applications, tutorials, and software built using the H2O open-source machine learning platform.
A collection of utilities and scripts for interactive data exploration, analysis, and automated modeling within Microsoft's Team Data Science Process.
A Julia package providing comprehensive clustering algorithms and validation metrics for data analysis.
A high-level machine learning library for Go with a Keras-like API, built on Gorgonia.
A lightweight Julia toolkit for working with time series data, providing efficient data structures and operations.
A Jupyter kernel that enables interactive computing with the Elixir programming language.
An intelligent data search and enrichment library for machine learning that automatically finds and adds relevant external features to ML pipelines.
A Rust library for creating directed hypergraphs where hyperedges can connect any number of vertices.
A DBI-compliant R interface to PostgreSQL, rewritten in C++ for improved performance and reliability.
An Elixir library for structured data extraction from websites, articles, and RSS/Atom feeds using information-retrieval techniques.
A scikit-learn compatible Python library for probabilistic regression, survival analysis, and probability distributions.
Automated infrastructure setup tool for training machine learning algorithms on AWS.
An OCaml kernel for Jupyter notebooks, providing an OCaml REPL with markdown/HTML documentation, LaTeX, and image embedding.
A scalable machine learning library that runs on Apache Hive, Spark, and Pig for distributed ML directly in SQL.
A browser extension that adds AI-powered code assistance to Jupyter Notebooks and Jupyter Lab using ChatGPT/GPT-4.
A JupyterLab extension to visualize CSV and JSON data interactively using Voyager 2.
An R data package providing an excerpt from Gapminder's global development data for teaching and examples.
A Node.js library implementing Support Vector Machines (SVM) for classification and regression tasks.
Open-Awesome is built by the community, for the community. Submit a project, suggest an awesome list, or help improve the catalog on GitHub.