Showing 7 of 7 projects
A Go tool and library for downloading URLs and files from Common Crawl and Wayback Machine web archives.
An Apache Spark framework for efficient data processing, extraction, and derivation from web archives and archival collections.
A collection of robust and fast Python tools for parsing, extracting, and analyzing web archive data, including a high-performance WARC parser.
A Node.js library for parsing and creating Web ARChive (WARC) files with support for Chrome, Puppeteer, and Electron.
A splitable Hadoop InputFormat for processing concatenated GZIP files and web archive (*.warc.gz) data efficiently in distributed systems.
A Node.js library for parsing CDXJ files produced by web archiving tools like Pywb.
A Hadoop/MapReduce tool that splits and partitions web archive records in (W)ARC files by MIME type and year.
Open-Awesome is built by the community, for the community. Submit a project, suggest an awesome list, or help improve the catalog on GitHub.