Showing 36 of 625 projects
A collection of slides and materials for talks on Jupyter, Vega, and data science tools.
A machine learning game that asks questions to guess objects using Bayesian inference and entropy minimization.
Python library for analyzing temporal networks and detecting dynamic communities.
A Python library for training hybrid recommendation systems using scikit-learn algorithms with zero code required.
Code for Kaggle's Accelerometer Biometric Competition, featuring scripts for data analysis and baseline predictions.
An R package that imports COBOL CopyBook data files directly into R as structured data frames.
A collection of network analysis helper functions for igraph in R, providing missing or niche utilities.
A set of Directus extensions for training, testing, and using machine learning models directly within Directus.
A Python command-line tool for encoding text sequences into vector-of-vector representations using word embeddings.
A Ruby extension for time series analysis with ARIMA, Kalman filtering, autocorrelation, and other statistical methods.
A Swift library for building predictions using linear regression with simple API and statistical insights.
Neovim plugin that launches a Jupyter/IPython console in a split to run files, lines, and Spyder‑IDE‑style runcells, with a ZMQ‑powered variable explorer and data viewer.
Pure Go library for reading and writing MATLAB .mat files (v5-v7.3+) without CGo dependencies.
A Visual Studio Code extension that enhances Quarto document editing with interactive code cells, Zotero citations, footnote highlighting, and inline code running.
A collection of pandas & scikit-learn compatible transformers for preprocessing and feature engineering 🛠
A Python library built on pandas and geopandas for linearly referenced data management, engineering, and analysis using the DataFrame accessor pattern (.lr).
An implementation of Dell Zhang's solution to Wikipedia's Participation Challenge on Kaggle.
A Python implementation of the clugen algorithm for generating multidimensional clusters with arbitrary distributions.
A Python library for data acquisition from I2C sensors on Raspberry Pi and MicroPython platforms.
A topic modeling project using Latent Dirichlet Allocation (LDA) to analyze and categorize Sarah Palin's released emails.
An R package providing easy access to 1,075+ complex network datasets from the Colorado Index of Complex Networks (ICON).
An R package for analyzing two-mode networks and extracting binary backbones from their one-mode projections.
Text and supporting code for Think Stats, 2nd Edition, a practical introduction to statistics for programmers.
A Julia implementation of the clugen algorithm for generating multidimensional clusters with arbitrary distributions.
Raku bindings for libsvm, providing support vector machine algorithms for classification, regression, and outlier detection.
Code for the Best Buy competition at Kaggle, focusing on mobile contest big data analysis.
A Python package for analyzing high-throughput single-cell imaging data, including protein abundance, endocytosis, and particle tracking.
A secure, scalable IPython Notebook service for managing multiple notebook servers with LDAP authentication and supervisor process management.
A VS Code extension for viewing large datasets (JSONL/Parquet/CSV) instantly without crashes, with 16 production LLM tokenizers for accurate token counting.
A no-code data visualization tool for creating interactive dashboards from CSV and Excel files.
Python package that projects customer retention rates and calculates lifetime value using shifted-beta-geometric distribution fitting.
A Julia library providing tools for working with tabular data, similar to pandas or R data frames.
A small dataset for the Best Buy mobile contest on Kaggle, containing product information and search queries.
A robust Rust library for regression analysis with sklearn-style estimators and full statistical inference.
A dependency-free statistics, linear algebra, and machine learning library for the V programming language, focused on product analytics.
A MATLAB/Octave implementation of the clugen algorithm for generating multidimensional clusters with arbitrary distributions.
Open-Awesome is built by the community, for the community. Submit a project, suggest an awesome list, or help improve the catalog on GitHub.