Open-Awesome
CategoriesAlternativesStacksSelf-HostedExplore
Open-Awesome

© 2026 Open-Awesome. Curated for the developer elite.

TermsPrivacyAboutGitHubRSS
  1. Home
  2. Web Archiving
  3. ArchiveSpark

ArchiveSpark

MITScalalatest-SNAPSHOT

An Apache Spark framework for efficient data processing, extraction, and derivation from web archives and archival collections.

GitHubGitHub
161 stars19 forks0 contributors

What is ArchiveSpark?

ArchiveSpark is an Apache Spark framework specifically designed for processing, extracting, and deriving data from web archives and other archival collections. It solves the problem of efficiently accessing and transforming raw archival data into more usable formats while maintaining data lineage. The framework enables researchers and developers to create derived datasets through filtering and extraction operations.

Target Audience

Data scientists, researchers, and developers working with web archives or archival collections who need to process, analyze, and extract value from large-scale historical data. This includes digital humanities researchers, web archivists, and data engineers working with temporal web data.

Value Proposition

Developers choose ArchiveSpark because it provides a specialized, efficient framework for archival data processing built on Apache Spark's distributed computing capabilities. Its unique modular architecture with customizable data specifications allows it to work with diverse archival collections while maintaining data lineage—a critical feature for reproducible research.

Overview

An Apache Spark framework for easy data processing, extraction as well as derivation for web archives and archival collections, developed at Internet Archive.

Use Cases

Best For

  • Processing and analyzing large-scale web archive collections
  • Extracting specific properties from archived web data for research
  • Creating derived datasets from archival collections with preserved lineage
  • Performing temporal analysis on historical web data
  • Generating hyperlink or knowledge graphs from archived web content
  • Transforming raw archival data into accessible formats like JSON

Not Ideal For

  • Real-time or streaming data processing applications
  • Small-scale data extraction tasks that don't require distributed computing
  • Projects needing extensive user interfaces or non-archival data formats
  • Teams without prior Apache Spark or distributed systems expertise

Pros & Cons

Pros

Efficient Distributed Processing

Leverages Apache Spark's distributed computing to handle large archival collections efficiently, as highlighted in its philosophy for scalable data access.

Data Lineage Tracking

Automatically reflects the lineage of derived values back to original sources, ensuring traceability for reproducible research and analysis.

Modular and Extensible

Customizable data specifications allow adaptation beyond web archives to any archival collection, supporting diverse data sources.

Specialized for Archival Workflows

Built-in support for formats like WARC/CDX and tools for temporal analysis make it ideal for web archive processing and derived corpus creation.

Cons

Steep Learning Curve

Requires deep knowledge of Apache Spark and distributed systems, making it challenging for newcomers without this background.

Internet Archive Dependency

Based on Sparkling, an internal library from Internet Archive, which may lead to vendor lock-in and limited control over future updates.

Niche Focus Limitations

Primarily designed for archival data, so it lacks features for general-purpose data processing or real-time applications, as admitted in its specialized use cases.

Frequently Asked Questions

Quick Stats

Stars161
Forks19
Contributors0
Open Issues4
Last commit9 months ago
CreatedSince 2015

Tags

#data-lineage#apache-spark#web-archives#spark#warc#internet-archive#webarchive#web-archiving#data-processing#temporal-analysis#distributed-computing#data-extraction

Built With

A
Apache Spark

Included in

Web Archiving2.5k
Auto-fetched 6 hours ago

Related Projects

Archives Unleashed ToolkitArchives Unleashed Toolkit

The Archives Unleashed Toolkit is an open-source toolkit for analyzing web archives.

Stars158
Forks33
Last commit7 months ago
Common Crawl Jupyter notebooksCommon Crawl Jupyter notebooks

Various Jupyter notebooks about Common Crawl data

Stars66
Forks11
Last commit21 days ago
Archives Unleashed NotebooksArchives Unleashed Notebooks

Various examples of notebooks for working with web archives with the Archives Unleashed Toolkit, and derivatives generated by the Archives Unleashed Toolkit.

Stars26
Forks5
Last commit3 years ago
Archives Research Compute HubArchives Research Compute Hub

Web application for distributed compute analysis of Archive-It web archive collections.

Stars20
Forks3
Last commit4 months ago
Community-curated · Updated weekly · 100% open source

Found a gem we're missing?

Open-Awesome is built by the community, for the community. Submit a project, suggest an awesome list, or help improve the catalog on GitHub.

Submit a projectStar on GitHub