Showing 3 of 3 projects
Scripts and tools to recreate the ELI5 dataset for long-form question answering research.
A Go tool and library for downloading URLs and files from Common Crawl and Wayback Machine web archives.
A collection of Jupyter notebooks for analyzing Common Crawl web archive data using columnar indexes and webgraph datasets.
Open-Awesome is built by the community, for the community. Submit a project, suggest an awesome list, or help improve the catalog on GitHub.