Showing 36 of 624 projects
A collection of R packages for interacting with Hadoop ecosystems, enabling big data analysis from R.
A Jupyter Notebook kernel for interactive data exploration and analysis using Apache Spark with Scala.
A high-performance, functional tabular data processing library for Clojure, similar to Python's Pandas or R's data.table.
A Nix-based framework for creating declarative and reproducible Jupyter environments with configurable kernels and extensions.
A collection of Jupyter notebooks providing examples and tutorials for the Bokeh interactive visualization library.
Easy pipelines for pandas DataFrames.
A tool to package, serve, and deploy any ML model on any platform using a GitOps approach.
A modern C++ toolkit for text retrieval and analysis, featuring indexing, ranking, topic modeling, classification, and language models.
A hyperparameter-free gradient boosting machine with a simple budget parameter, built for high performance with Rust and bindings for Python and R.
IPython-based environment for reproducible machine learning research with unified wrappers for multiple ML libraries.
A Python package for stacking (stacked generalization) with both functional and scikit-learn compatible APIs.
A quick reference guide to the most commonly used patterns and functions in PySpark SQL.
A curated list of proven AI use cases that generate business value across departments and industries.
ADO.NET provider and native bindings for DuckDB, enabling C# applications to interact with the in-process analytical database.
Capture, analyze, and transform messy Jupyter notebooks into production data pipelines with just two lines of code.
A curated guide to essential R packages organized by their role in the data science workflow.
A Python library for comparing Pandas, Polars, Spark, and Snowpark DataFrames with detailed reporting and flexible matching.
A curated collection of free resources to help deepen your understanding of the R programming language.
An R package providing a lightweight frontend to use Apache Spark for distributed data processing from R.
A Julia package for fitting linear and generalized linear models with comprehensive statistical functionality.
A comprehensive roadmap chart and resource guide for aspiring data scientists, based on insights from Silicon Valley tech companies.
An R package that simplifies data import and export by automatically selecting the correct function based on file extension.
A tutorial series comparing how to implement data science concepts and build applications in both Python and R ecosystems.
A VS Code extension for visually exploring, cleaning, and transforming tabular data with automatic Pandas code generation.
A deep learning system that classifies food images into 230 categories and retrieves matching recipes using convolutional neural networks.
A Python library providing evaluation metrics and diagnostic tools for recommender systems.
A simplified Keras-like framework for PyTorch that reduces boilerplate code for training neural networks.
An optimized distributed gradient boosting library for fast and accurate machine learning on large datasets.
A collection of IPython notebooks containing machine learning experiments and examples using scikit-learn and related Python libraries.
A Python library for generating high-quality synthetic tabular data using GANs, diffusion models, and large language models.
A tidy API for graph manipulation in R, providing dplyr verbs and igraph algorithms for network analysis.
A modern, high-performance technical analysis library built in Rust with Python and WebAssembly bindings.
A machine learning integrations library for TypeDB, enabling graph algorithms and Graph Neural Networks on strongly-typed graph data.
An R package that automates exploratory data analysis and data treatment with one-line reports and visualizations.
An engine for ML/data tracking, visualization, explainability, drift detection, and dashboards, integrated with Polyaxon.
A Neovim plugin providing language support, code execution, and preview features for working with Quarto documents.
Open-Awesome is built by the community, for the community. Submit a project, suggest an awesome list, or help improve the catalog on GitHub.