Showing 36 of 624 projects
Python library for analyzing temporal networks and detecting dynamic communities.
A Python library for training hybrid recommendation systems using scikit-learn algorithms with zero code required.
A machine learning game that asks questions to guess objects using Bayesian inference and entropy minimization.
Code for Kaggle's Accelerometer Biometric Competition, featuring scripts for data analysis and baseline predictions.
A collection of network analysis helper functions for igraph in R, providing missing or niche utilities.
An R package that imports COBOL CopyBook data files directly into R as structured data frames.
A Ruby extension for time series analysis with ARIMA, Kalman filtering, autocorrelation, and other statistical methods.
A Python command-line tool for encoding text sequences into vector-of-vector representations using word embeddings.
A set of Directus extensions for training, testing, and using machine learning models directly within Directus.
A Swift library for building predictions using linear regression with simple API and statistical insights.
Neovim plugin that launches a Jupyter/IPython console in a split to run files, lines, and Spyder‑IDE‑style runcells, with a ZMQ‑powered variable explorer and data viewer.
Pure Go library for reading and writing MATLAB .mat files (v5-v7.3+) without CGo dependencies.
A Visual Studio Code extension that enhances Quarto document editing with interactive code cells, Zotero citations, footnote highlighting, and inline code running.
A collection of pandas & scikit-learn compatible transformers for preprocessing and feature engineering 🛠
An implementation of Dell Zhang's solution to Wikipedia's Participation Challenge on Kaggle.
A Python library built on pandas and geopandas for linearly referenced data management, engineering, and analysis using the DataFrame accessor pattern (.lr).
A Python implementation of the clugen algorithm for generating multidimensional clusters with arbitrary distributions.
A Python library for data acquisition from I2C sensors on Raspberry Pi and MicroPython platforms.
An R package for analyzing two-mode networks and extracting binary backbones from their one-mode projections.
A topic modeling project using Latent Dirichlet Allocation (LDA) to analyze and categorize Sarah Palin's released emails.
An R package providing easy access to 1,075+ complex network datasets from the Colorado Index of Complex Networks (ICON).
Code for the Best Buy competition at Kaggle, focusing on mobile contest big data analysis.
A Julia implementation of the clugen algorithm for generating multidimensional clusters with arbitrary distributions.
Text and supporting code for Think Stats, 2nd Edition, a practical introduction to statistics for programmers.
Raku bindings for libsvm, providing support vector machine algorithms for classification, regression, and outlier detection.
A Python package for analyzing high-throughput single-cell imaging data, including protein abundance, endocytosis, and particle tracking.
Python package that projects customer retention rates and calculates lifetime value using shifted-beta-geometric distribution fitting.
A no-code data visualization tool for creating interactive dashboards from CSV and Excel files.
A VS Code extension for viewing large datasets (JSONL/Parquet/CSV) instantly without crashes, with 16 production LLM tokenizers for accurate token counting.
A secure, scalable IPython Notebook service for managing multiple notebook servers with LDAP authentication and supervisor process management.
A small dataset for the Best Buy mobile contest on Kaggle, containing product information and search queries.
A Julia library providing tools for working with tabular data, similar to pandas or R data frames.
A robust Rust library for regression analysis with sklearn-style estimators and full statistical inference.
A dependency-free statistics, linear algebra, and machine learning library for the V programming language, focused on product analytics.
A MATLAB/Octave implementation of the clugen algorithm for generating multidimensional clusters with arbitrary distributions.
Structured U.S. drinking water quality data for 41,000+ ZIP codes, including EPA violations, lead/copper levels, PFAS, radon, flood risk, and Home Safety Scores.
Open-Awesome is built by the community, for the community. Submit a project, suggest an awesome list, or help improve the catalog on GitHub.