Showing 31 of 31 projects
An open-source, extensible, web-scale, archival-quality web crawler from the Internet Archive.
An open-source, extensible, web-scale, archival-quality web crawler from the Internet Archive.
A preconfigured web crawler for backing up websites, producing WARC files with a live dashboard and dynamic ignore patterns.
A browser extension and desktop app for interactive, high-fidelity web archiving directly in the browser.
A standalone Docker container for high-fidelity, browser-based web archiving crawls using Puppeteer and Brave.
A distributed and persistent web archive replay system that uses IPFS to store and serve WARC files.
Streaming WARC/ARC library for fast web archive IO
WarcDB is an SQLite-based file format that makes web crawl data easier to share and query.
A graphical desktop application that simplifies web archiving by providing a one-click interface to preserve and replay web pages using Heritrix and OpenWayback.
A high-fidelity, browser-based web archiving library and CLI for capturing single web pages with provenance.
A high-fidelity, user-scriptable archival web crawler using Chrome/Chromium to preserve JavaScript-rendered content.
An Apache Spark framework for efficient data processing, extraction, and derivation from web archives and archival collections.
A web application for searching, browsing, and analyzing archived web content (ARC/WARC files) with a Solr backend.
A collection of robust and fast Python tools for parsing, extracting, and analyzing web archive data, including a high-performance WARC parser.
A Node.js library for parsing and creating Web ARChive (WARC) files with support for Chrome, Puppeteer, and Electron.
An offline-first web browser that archives, searches, and crawls websites for personal use.
A command-line tool for archiving web pages into WARC files and replaying them locally.
A Rust library for reading and writing WARC (Web ARChive) files.
A Rails engine for discovering web archives in WARC and ARC formats with faceted search and advanced discovery options.
A Python tutorial demonstrating how to access and process Common Crawl's web archive datasets (WARC, WET, WAT) using tools like warcio, cdx_toolkit, and DuckDB.
Web archiving using Google Chrome
A command-line tool and Rust library for handling Web ARChive (WARC) files.
A customizable Scala crawler for creating personal web archives in WARC/CDX format.
A Python tool that scans WARC web archive files for viruses and NSFW content using AI and antivirus detection.
A BadgerDB-based server for indexing and serving WARC file contents with CDX capture index support.
A Scala/Spark library for efficient processing, extraction, and derivation of web archive data (CDX/WARC).
A Python CLI tool for extracting and validating WARC and WACZ web archive files.
Java-based web archive deduplication tool that identifies duplicates and converts them to reference records in WARC files.
A splitable Hadoop InputFormat for processing concatenated GZIP files and web archive (*.warc.gz) data efficiently in distributed systems.
A command-line tool for performing various gzip, ARC, WARC, and XML tasks on web archive files.
A Hadoop/MapReduce tool that splits and partitions web archive records in (W)ARC files by MIME type and year.
Open-Awesome is built by the community, for the community. Submit a project, suggest an awesome list, or help improve the catalog on GitHub.