Showing 36 of 80 projects
An offline-first web browser that archives, searches, and crawls websites for personal use.
A portable concurrent Memento aggregator CLI and server for retrieving archived web pages from multiple sources.
A Python toolkit for extracting, filtering, and analyzing data from web archives, JSON files, and imageboards.
A command-line tool for archiving web pages into WARC files and replaying them locally.
A dockerized, queued web archiver using Chrome headless to create high-fidelity WARC files from URLs.
A Rust library for reading and writing WARC (Web ARChive) files.
A Java library for reading and writing WARC files with a typed, extensible API and high-performance NIO-based parsing.
A Chrome extension that integrates live web browsing with archived copies using the Memento protocol.
Converts WARC web archive files to static HTML with relative links for offline browsing or rehosting.
MapReduce tools for bulk indexing of web archive WARC/ARC files into ZipNum sharded CDX clusters on Hadoop, EMR, or local systems.
A Python tutorial demonstrating how to access and process Common Crawl's web archive datasets (WARC, WET, WAT) using tools like warcio, cdx_toolkit, and DuckDB.
A high-performance, RocksDB-based capture index server for web archives, supporting OpenWayback and PyWb protocols.
A collection of salvaged websites, articles, contributions, text, and documentation preserved from various sources.
A collection of salvaged websites, articles, contributions, text, and documentation preserved from the web.
Converts HTTrack website crawls into standardized WARC files for web archiving and preservation.
A command-line tool and Rust library for handling Web ARChive (WARC) files.
A personal web archive and search system that runs in a Docker container, allowing you to browse and search archived web pages.
A customizable Scala crawler for creating personal web archives in WARC/CDX format.
A Python script that converts offline web resources into a single WARC file for archiving.
A comprehensive open educational resource for teaching web archiving concepts, practices, and tools.
A resilient and configurable Python tool for exploring, analyzing, transforming, and extracting data from WARC (Web ARChive) files.
A DuckDB extension to query web archive CDX APIs (Wayback Machine & Common Crawl) directly from SQL with smart query pushdown.
A Go library for reading and parsing WARC and ARC web archive formats with specialized utilities for web archiving workflows.
A job server for distributed compute analysis of web archive (WARC) collections.
A Python tool that scans WARC web archive files for viruses and NSFW content using AI and antivirus detection.
A Scala/Spark library for efficient processing, extraction, and derivation of web archive data (CDX/WARC).
A BadgerDB-based server for indexing and serving WARC file contents with CDX capture index support.
A Python command-line client for downloading web archive (WARC) files from Archive-It and Webrecorder WASAPI Data Transfer APIs.
A Python CLI tool for extracting and validating WARC and WACZ web archive files.
A command-line tool to playback archived webpages from the Wayback Machine using GitHub as a source.
A command-line tool and Go package for archiving webpages to IPFS with support for local nodes and remote pinning services.
A framework for profiling web archives to summarize their holdings using compact SURT-based maps.
Extracts hyperlinks from files using Apache Tika for batch processing and web archiving workflows.
A virtual machine and walkthrough for setting up and using the Heritrix web crawler for web archiving.
Java-based web archive deduplication tool that identifies duplicates and converts them to reference records in WARC files.
A CLI tool to test URL availability and retrieve Internet Archive snapshots, with output in JSON, CSV, or BoltDB.
Open-Awesome is built by the community, for the community. Submit a project, suggest an awesome list, or help improve the catalog on GitHub.