Showing 36 of 624 projects
CLI task management & automation tool
An embeddable C++ storage engine for dense and sparse multi-dimensional arrays, dataframes, and key-value stores.
A Python framework and Rust-based distributed processing engine for stateful event and stream processing.
A Julia machine learning framework providing a unified interface and meta-algorithms for over 200 models.
Automatically visualize any dataset with a single line of code, including data quality assessment and fixes.
A minimal benchmark comparing scalability, speed, and accuracy of popular open-source machine learning libraries for binary classification.
A Python library for probabilistic prediction using natural gradient boosting, built on scikit-learn.
A curated list of awesome Apache Spark packages, libraries, and resources for data engineers and scientists.
A high-performance Python package for fast, multi-threaded manipulation of large tabular datasets, inspired by R's data.table.
An R package for estimating causal effects in time series using Bayesian structural time-series models.
A Python library for exploratory analysis, diagnostics, and visualization of Bayesian models.
A flexible and fast package for in-memory tabular data manipulation and analysis in the Julia programming language.
A curated collection of academic papers on data mining and machine learning techniques for fraud detection across various domains.
A meta-package for installing and loading core R packages for data science that share common design principles.
A collection of R packages for data science that share common design principles and work together seamlessly.
An open-source Python toolkit providing a comprehensive collection of algorithms for interpreting and explaining machine learning models and datasets.
Create blogs and websites with R Markdown, integrating dynamic R code, graphics, and technical writing elements.
A lightweight MongoDB schema analyzer that reveals document structure, field frequencies, and data outliers.
A comprehensive R package that embeds Python within R sessions, enabling seamless interoperability between the two languages.
Python code and examples for Bayesian statistics from the book 'Think Bayes: Bayesian Statistics Made Simple'.
A curated collection of 500+ resources for data analysis and data science, covering Python, SQL, ML, visualization, roadmaps, and interview prep.
A native R kernel for Jupyter notebooks, enabling R programming within the Jupyter ecosystem.
A unified interface and infrastructure for machine learning in R, supporting classification, regression, clustering, and survival analysis.
Automated machine learning library for production and analytics, handling feature engineering, model selection, and hyperparameter optimization.
Hyperopt-sklearn automates hyperparameter optimization and model selection for scikit-learn machine learning pipelines.
Machine learning with dataframes
Python implementation of the Boruta all-relevant feature selection method with scikit-learn compatibility.
A curated collection of 60 ChatGPT prompts for data science tasks, from model building to code explanation.
A satirical programming language designed to mock enterprise software development culture with intentionally cumbersome syntax and corporate jargon.
A Go machine learning library with online learning capabilities and a variety of implemented models.
A curated reading list and syllabus for a Stanford discussion class on applied data science topics.
A JupyterLab extension for version control using Git, enabling Git operations directly within the JupyterLab interface.
A Python package for concise, transparent, and accurate predictive modeling with sklearn-compatible interpretable models.
A deprecated repository for community-contributed Keras extensions like layers, activations, and loss functions.
An open-source Python repository providing around 40 feature selection algorithms for machine learning applications.
A Python library that automatically extracts schema, statistics, and sensitive entities (PII/NPI) from datasets.
Open-Awesome is built by the community, for the community. Submit a project, suggest an awesome list, or help improve the catalog on GitHub.