Open-Awesome
CategoriesAlternativesStacksSelf-HostedExplore
Open-Awesome

© 2026 Open-Awesome. Curated for the developer elite.

TermsPrivacyAboutGitHubRSS
  1. Home
  2. Web Archiving
  3. webarchive-discovery

webarchive-discovery

Javawarc-discovery-3.3.1

A toolkit for indexing and exploring web archive content from ARC and WARC files using OpenSearch/Elasticsearch.

Visit WebsiteGitHubGitHub
133 stars26 forks0 contributors

What is webarchive-discovery?

Web Archive Discovery is a toolkit for indexing and searching web archive content stored in ARC and WARC files. It extracts data from these archives and indexes it into search engines like OpenSearch or Elasticsearch, enabling users to explore and discover historical web materials. The project solves the problem of making large-scale web archives searchable and accessible for research or archival purposes.

Target Audience

Digital archivists, librarians, researchers, and institutions managing web archives who need to build searchable indexes of archived web content.

Value Proposition

Developers choose Web Archive Discovery for its specialized focus on web archive indexing, integration with modern search engines, and self-hosted deployment options, providing a flexible and scalable solution for exploring archived web data.

Overview

Please note that the warc-indexer tool & code is now supported by NetArchiveSuite. The 'warc-indexer' directory and code that exists in this repo is now only for reference. For support and issues of 'warc-indexer', please communicate with NetArchiveSuite.

Use Cases

Best For

  • Indexing large collections of WARC or ARC files for search
  • Building a searchable web archive for research institutions
  • Self-hosting a web archive discovery platform
  • Integrating web archive content with OpenSearch or Elasticsearch
  • Developing tools for digital preservation and archival access
  • Creating explorable interfaces for historical web data

Not Ideal For

  • Projects requiring real-time indexing of live web content or streaming data
  • Teams seeking a fully managed, cloud-based archive search solution with minimal setup
  • Small-scale personal archives where deploying and maintaining a search engine is overly complex
  • Developers unfamiliar with Java, Docker, or command-line tools for system administration

Pros & Cons

Pros

Specialized Web Archive Focus

Tailored specifically for ARC and WARC files, offering optimized extraction and indexing for historical web data, as highlighted in its core purpose of making archive contents explorable.

Modern Search Engine Integration

Integrates seamlessly with OpenSearch and Elasticsearch, enabling scalable search capabilities, and includes a Docker Compose setup for easy local development and testing.

Solr Schema Portability

Ports Solr schemas to OpenSearch with minimal adjustments, easing migration for users transitioning from Solr-based systems, as noted in the README's compatibility details.

Command-Line Flexibility

Provides a Java-based CLI tool for batch indexing WARC files, allowing automation and customization in large-scale archive processing workflows.

Cons

Split Maintenance Model

The core warc-indexer tool is now supported by NetArchiveSuite, making this repository potentially outdated for active development and leading to confusion in issue tracking.

Complex Initial Setup

Requires configuring Docker, OpenSearch/Elasticsearch, and Java environments, which can be daunting for users without prior experience in these technologies.

Sparse Documentation

Relies on a wiki for documentation, which may lack detailed guides or updates, as evidenced by the minimal instructions in the README for advanced use cases.

Frequently Asked Questions

Quick Stats

Stars133
Forks26
Contributors0
Open Issues90
Last commit8 months ago
CreatedSince 2012

Tags

#digital-preservation#java#warc-indexing#docker#search-engine#web-archiving#opensearch#elasticsearch#data-mining

Built With

E
Elasticsearch
M
Maven
J
Java
D
Docker
O
OpenSearch

Links & Resources

Website

Included in

Web Archiving2.5k
Auto-fetched 5 hours ago

Related Projects

hyphehyphe

Websites crawler with built-in exploration and control web interface

Stars384
Forks62
Last commit2 months ago
SolrWaybackSolrWayback

A search interface and wayback machine for the UKWA Solr based warc-indexer framework.

Stars145
Forks28
Last commit1 day ago
herehere

Please note that the warc-indexer tool & code is now supported by NetArchiveSuite. The 'warc-indexer' directory and code that exists in this repo is now only for reference. For support and issues of 'warc-indexer', please communicate with NetArchiveSuite.

Stars133
Forks26
Last commit8 months ago
MinkMink

Chrome extension that uses Memento to indicate that a page a user is viewing on the live web has an archived copy and to give the user access to the copy

Stars59
Forks3
Last commit11 months ago
Community-curated · Updated weekly · 100% open source

Found a gem we're missing?

Open-Awesome is built by the community, for the community. Submit a project, suggest an awesome list, or help improve the catalog on GitHub.

Submit a projectStar on GitHub