Showing 10 of 10 projects
A curated list of resources, tools, and services for web archiving, from acquisition and replay to analysis and community.
A curated list of resources, tools, and services for web archiving, from acquisition and replay to analysis and community.
A high-fidelity, user-scriptable archival web crawler using Chrome/Chromium to preserve JavaScript-rendered content.
A curated list of software, literature, and resources for the Memento protocol (RFC7089) enabling time-based access to archived web content.
A Node.js library for parsing and creating Web ARChive (WARC) files with support for Chrome, Puppeteer, and Electron.
A collection of Jupyter notebooks for analyzing Common Crawl web archive data using columnar indexes and webgraph datasets.
A dockerized, queued web archiver using Chrome headless to create high-fidelity WARC files from URLs.
A Python tool that scans WARC web archive files for viruses and NSFW content using AI and antivirus detection.
Extracts hyperlinks from files using Apache Tika for batch processing and web archiving workflows.
A Node.js library for parsing CDXJ files produced by web archiving tools like Pywb.
Open-Awesome is built by the community, for the community. Submit a project, suggest an awesome list, or help improve the catalog on GitHub.