Showing 21 of 57 projects
A Rails engine for discovering web archives in WARC and ARC formats with faceted search and advanced discovery options.
Converts HTTrack website crawls into standardized WARC files for web archiving and preservation.
A customizable Scala crawler for creating personal web archives in WARC/CDX format.
A Python script that converts offline web resources into a single WARC file for archiving.
A comprehensive open educational resource for teaching web archiving concepts, practices, and tools.
A resilient and configurable Python tool for exploring, analyzing, transforming, and extracting data from WARC (Web ARChive) files.
A Go library for reading and parsing WARC and ARC web archive formats with specialized utilities for web archiving workflows.
A Wagtail-based CMS and Ansible playbooks for managing the Localore: Finding America project.
A Python tool that scans WARC web archive files for viruses and NSFW content using AI and antivirus detection.
A BadgerDB-based server for indexing and serving WARC file contents with CDX capture index support.
A Python command-line client for downloading web archive (WARC) files from Archive-It and Webrecorder WASAPI Data Transfer APIs.
An institutional repository application for managing, preserving, and providing access to scholarly research outputs.
A framework for profiling web archives to summarize their holdings using compact SURT-based maps.
Extracts hyperlinks from files using Apache Tika for batch processing and web archiving workflows.
A CLI tool to test URL availability and retrieve Internet Archive snapshots, with output in JSON, CSV, or BoltDB.
Java-based web archive deduplication tool that identifies duplicates and converts them to reference records in WARC files.
A virtual machine and walkthrough for setting up and using the Heritrix web crawler for web archiving.
A Java command-line application for downloading web archive (WARC) files from the WASAPI (Web Archiving Systems API).
A command-line tool for performing various gzip, ARC, WARC, and XML tasks on web archive files.
Converts bag-nabit datasets from ZIP archives into full-content WARC files for web archiving.
Intelligent web page comparison tool with Wayback Machine artifact cleaning, visual regression testing, and significance scoring.
Open-Awesome is built by the community, for the community. Submit a project, suggest an awesome list, or help improve the catalog on GitHub.