Showing 31 of 103 projects
A benchmark dataset and meta self-learning method for multi-source domain adaptation in scene text recognition.
A YOLO-based object detection system specifically trained to identify DJI drones in images and video.
A multi-sensor dataset for autonomous vehicle and robot navigation, featuring synchronized camera, LiDAR, IMU, and GNSS data collected in urban environments.
A public dataset of field images with segmentation masks and plant type annotations for computer vision in precision agriculture.
An R package providing 2,260 network datasets in igraph format from diverse sources like social networks, animal interactions, and movie co-stars.
A symbolic programming library built on JAX for concise, explicit, and optimized machine learning computations.
A weekly updated dataset of Dungeons & Dragons characters submitted to character sheet web applications, with over 7,900 entries and standardized fields.
A Python devkit for working with the Boreas and Boreas Road Trip all-weather autonomous driving datasets.
A dataset of NBA game summaries aligned with box- and line-scores for data-to-text generation research.
Replication package and dataset for a research paper on software architecture practices in ROS-based robotic systems.
A characteristic-rich dataset for factoid question answering with explicit question specifications to enable fine-grained QA system evaluation.
An open dataset for learning-based temporal analysis of PE malware, containing over 130,000 Windows PE files with feature vectors and metadata.
A city-scale dataset and platform for learning holistic 3D structures from panoramic and perspective imagery with detailed annotations.
A research project investigating how packers affect the accuracy of static machine-learning malware classifiers.
A package manager for machine learning datasets and models with a CLI and self-hostable registry.
An enriched dataset for Natural Language Generation research, providing intermediate representations for pipeline tasks like lexicalization and aggregation.
A Delphi component that enables SQL queries on TDataSet descendants with its own parser and engine, no DLL required.
A repository of extracted game data for Brave Frontier, including units, items, skills, and missions across Global, JP, and EU servers.
A curated dataset of packed and unpacked PE executables for training machine learning models to detect packing.
A curated collection of resources, datasets, and methods for 3D LiDAR-based Moving Object Segmentation (MOS) research.
Presentation slides from the useR 2022 conference talk about the untold story of the palmerpenguins dataset.
A collection of Danish wordlists for password cracking and security testing.
A dataset for context-aware natural language generation in task-oriented spoken dialogue systems for public transport information.
A dataset of ELF files packed with various packers for training machine learning models on executable packing detection.
A dataset of PE files packed with various packers for malware analysis and security research.
A multi-extract, multi-level dataset of Mozilla Bugzilla issue tracking history spanning 15 years for software engineering research.
A small dataset for the Best Buy mobile contest on Kaggle, containing product information and search queries.
A dataset of 9,188 OCL expressions from 504 EMF meta-models in 245 GitHub repositories for empirical OCL studies.
Structured U.S. drinking water quality data for 41,000+ ZIP codes, including EPA violations, lead/copper levels, PFAS, radon, flood risk, and Home Safety Scores.
A Python library for handling datasets in an XSLX-based format with conversion to ARFF, CSV, and FilelessDataset structures.
A Turkish news archive containing NAYN.CO articles and metadata for research and analysis.
Open-Awesome is built by the community, for the community. Submit a project, suggest an awesome list, or help improve the catalog on GitHub.