Showing 24 of 60 projects
A Python library providing comprehensive metrics for fair and thorough evaluation of recommender systems.
A high-performance, large-scale statistical machine learning library written in Common Lisp.
A CPU and GPU-accelerated matrix library optimized for high-performance data mining operations.
Course materials for GWU's Data Mining and Machine Learning classes covering preprocessing, modeling, and practical Kaggle applications.
A Go library implementing essential machine learning algorithms including linear regression, logistic regression, and neural networks.
A Python machine learning and informatics suite for analyzing, mining, and modeling chemical and materials data.
A multi-platform data-mining and visualization library for RAD Studio, supporting in-memory databases, pivot tables, and big data.
A toolkit for indexing and exploring web archive content from ARC and WARC files using OpenSearch/Elasticsearch.
A toolkit for indexing and exploring web archive content from ARC and WARC files using OpenSearch/Elasticsearch.
A collection of libraries for large-scale data processing in Hadoop ecosystems, including Spark, Pig, and incremental MapReduce.
A Julia package providing high-performance, configurable tokenizers and sentence splitters for natural language processing.
A comprehensive Java library for statistics, data mining, and machine learning with interactive notebook support.
A JRuby gem providing Ruby interfaces for Weka's machine learning and data mining algorithms.
A Go implementation of k-modes and k-prototypes clustering algorithms for categorical and mixed data.
A framework and GUI wizard for extracting structured information from tables in scientific literature, particularly biomedical publications.
Generates text using Markov chains trained on 4chan board data via command-line tool.
A fast high-level web crawling and scraping framework for Elixir, built on Broadway.
A workflow engine that unifies feature engineering and machine learning using a column-oriented data processing paradigm.
A K-Means clustering algorithm implementation for iOS with multi-dimensional clustering, data mining, and image compression support.
Elixir implementation of the CLOPE algorithm for fast clustering of transactional data.
An end-to-end Python system for time series data analysis with in-database machine learning using TDengine.
An R package that imports COBOL CopyBook data files directly into R as structured data frames.
Elixir implementation of the ROCK clustering algorithm for categorical data.
An iOS implementation of the Fuzzy C-Means clustering algorithm for machine learning tasks like data mining and image compression.
Open-Awesome is built by the community, for the community. Submit a project, suggest an awesome list, or help improve the catalog on GitHub.