Showing 36 of 625 projects
A Python library that brings R's dplyr data manipulation syntax to pandas DataFrames using a pipe operator.
A high-performance, functional tabular data processing library for Clojure, similar to Python's Pandas or R's data.table.
A collection of R packages for interacting with Hadoop ecosystems, enabling big data analysis from R.
A Jupyter Notebook kernel for interactive data exploration and analysis using Apache Spark with Scala.
A Nix-based framework for creating declarative and reproducible Jupyter environments with configurable kernels and extensions.
A collection of Jupyter notebooks providing examples and tutorials for the Bokeh interactive visualization library.
Easy pipelines for pandas DataFrames.
A tool to package, serve, and deploy any ML model on any platform using a GitOps approach.
A modern C++ toolkit for text retrieval and analysis, featuring indexing, ranking, topic modeling, classification, and language models.
A hyperparameter-free gradient boosting machine with a simple budget parameter, built for high performance with Rust and bindings for Python and R.
IPython-based environment for reproducible machine learning research with unified wrappers for multiple ML libraries.
A quick reference guide to the most commonly used patterns and functions in PySpark SQL.
A Python package for stacking (stacked generalization) with both functional and scikit-learn compatible APIs.
ADO.NET provider and native bindings for DuckDB, enabling C# applications to interact with the in-process analytical database.
A curated list of proven AI use cases that generate business value across departments and industries.
Capture, analyze, and transform messy Jupyter notebooks into production data pipelines with just two lines of code.
A curated guide to essential R packages organized by their role in the data science workflow.
A Python library for comparing Pandas, Polars, Spark, and Snowpark DataFrames with detailed reporting and flexible matching.
A curated collection of free resources to help deepen your understanding of the R programming language.
An R package providing a lightweight frontend to use Apache Spark for distributed data processing from R.
A Julia package for fitting linear and generalized linear models with comprehensive statistical functionality.
A comprehensive roadmap chart and resource guide for aspiring data scientists, based on insights from Silicon Valley tech companies.
An R package that simplifies data import and export by automatically selecting the correct function based on file extension.
A tutorial series comparing how to implement data science concepts and build applications in both Python and R ecosystems.
A VS Code extension for visually exploring, cleaning, and transforming tabular data with automatic Pandas code generation.
A deep learning system that classifies food images into 230 categories and retrieves matching recipes using convolutional neural networks.
An optimized distributed gradient boosting library for fast and accurate machine learning on large datasets.
A Python library providing evaluation metrics and diagnostic tools for recommender systems.
A simplified Keras-like framework for PyTorch that reduces boilerplate code for training neural networks.
A collection of IPython notebooks containing machine learning experiments and examples using scikit-learn and related Python libraries.
A Python library for generating high-quality synthetic tabular data using GANs, diffusion models, and large language models.
A tidy API for graph manipulation in R, providing dplyr verbs and igraph algorithms for network analysis.
A modern, high-performance technical analysis library built in Rust with Python and WebAssembly bindings.
A machine learning integrations library for TypeDB, enabling graph algorithms and Graph Neural Networks on strongly-typed graph data.
An R package that automates exploratory data analysis and data treatment with one-line reports and visualizations.
An engine for ML/data tracking, visualization, explainability, drift detection, and dashboards, integrated with Polyaxon.
Open-Awesome is built by the community, for the community. Submit a project, suggest an awesome list, or help improve the catalog on GitHub.