Showing 36 of 624 projects
An open-source machine learning system for the end-to-end data science lifecycle from data preparation to model serving.
An open-source Java framework for rapid development of machine learning and statistical applications with large dataset support.
A modern, object-oriented machine learning framework for R, providing efficient building blocks for ML workflows.
A Go kernel for Jupyter notebooks that compiles each cell for fast execution and full Go compatibility.
A Python library for building Generalized Additive Models (GAMs) with a scikit-learn-like API, emphasizing interpretability and performance.
A Python package providing specialized statistical algorithms for graph and network analysis.
A meta gem that bundles scientific computing and visualization libraries for Ruby, enabling data analysis and plotting.
A PyTorch-based framework for training and validating models that produce high-quality embeddings for metric learning and retrieval tasks.
An R interface for Apache Spark that enables distributed data processing, machine learning, and SQL queries using familiar R syntax.
A scikit-learn compatible Python module for multi-label classification tasks.
A fast, ergonomic machine learning library for Rust with broad algorithm coverage and WASM-first defaults.
A strongly-typed Scala API for TensorFlow, providing functionality similar to the official Python API with additional features.
A Ruby kernel for Jupyter notebooks, enabling interactive data science and computational workflows in Ruby.
Learn statistics through Python with real-world examples like analyzing marijuana price data across US states.
A Ruby machine learning library with a Scikit-Learn-like interface for classification, regression, clustering, and dimensionality reduction.
A curated collection of resources for Go-based data analysis, visualization, machine learning, and data science.
An R package for the quantitative analysis of textual data, providing comprehensive tools for natural language processing and text management.
An automated feature generation framework for tabular data that discovers expert-level features to boost machine learning model performance.
A curated list of awesome cheminformatics software, libraries, resources, and tools, primarily command-line based and open-source.
A curated list of resources for R Shiny, including tutorials, packages, deployment guides, and app examples.
A Jupyter kernel for Clojure, enabling Clojure code execution in Jupyter Lab, Notebook, and Console.
Convert IPython/Jupyter notebooks to markdown and back, enabling seamless editing of notebooks as markdown files.
A JupyterLab extension that integrates GPT-4 as a code interpreter, translating natural language to Python and executing it automatically.
A comprehensive statistical computation library for Rust, providing distributions, functions, and utilities for scientific computing.
An open-source image analysis software package for plant phenotyping using computer vision.
A biomedical knowledge graph integrating 20 resources to describe 17,080 diseases with over 4 million relationships across ten biological scales.
A pure Java machine learning library with no external dependencies, offering a wide collection of algorithms and parallel execution support.
A modular deep learning framework for PyTorch to build neural networks on heterogeneous tabular data.
An R package providing comprehensive historical soccer match datasets and analysis functions for European and MLS leagues.
A Python library for introductory data science education, developed for Berkeley's Data 8 course.
An R binding package for calling Google Earth Engine API from within R, integrating with the R spatial ecosystem.
A Neovim plugin that provides real-time, bidirectional synchronization with Jupyter Notebook using Selenium automation.
An open-source toolkit for auditing bias and experimenting with fairness methods in machine learning models.
:bowtie: Create a dashboard with python!
An R package for detecting statistically significant breakpoints in time series using robust energy statistics.
A Python library that brings R's dplyr data manipulation syntax to pandas DataFrames using a pipe operator.
Open-Awesome is built by the community, for the community. Submit a project, suggest an awesome list, or help improve the catalog on GitHub.