Open-Awesome
CategoriesAlternativesStacksSelf-HostedExplore
Open-Awesome

© 2026 Open-Awesome. Curated for the developer elite.

TermsPrivacyAboutGitHubRSS
  1. Home
  2. Tags
  3. Web Scraping

Web Scraping

190 projects

Showing 36 of 190 projects

Crawler
CrawlerElixir

A high-performance web crawler and scraper built in Elixir with worker pooling and rate limiting.

#elixir#spider#offline
Stars957
Forks89
Last commit1 month ago
xidel
xidelPascal

A command-line tool to extract data from HTML/XML pages and JSON APIs using CSS, XPath, XQuery, JSONiq, and pattern matching.

#rest#css-selectors#http
Stars840
Forks46
Last commit1 year ago
Spidr
SpidrRuby

A versatile Ruby web spidering library for crawling sites, domains, or specific links with extensive filtering and callback support.

#web-crawling#spider#crawler
Stars835
Forks107
Last commit6 months ago
ReadabilityKit
ReadabilityKitSwift

A Swift library for extracting article previews including title, description, images, and metadata from web pages.

#content-parsing#ios#metadata-extraction
Stars835
Forks79
Last commit3 years ago
jvppeteer
jvppeteerJava

A Java API for controlling Chrome and Firefox browsers via DevTools and WebDriver-bidi protocols.

#chrome#puppeteer#screenshot
Stars808
Forks170
Last commit12 days ago
suckit
suckitRust

A Rust-based command-line tool for recursively downloading entire websites for offline browsing.

#hacktoberfest#archiving#webscraping
Stars806
Forks44
Last commit4 months ago
cdp
cdpGo

Type-safe Go bindings for the Chrome DevTools Protocol, enabling browser automation and debugging.

#rpc-client#cdp#devtools-protocol
Stars794
Forks50
Last commit7 months ago
htmlquery
htmlqueryGo

A Go package for querying HTML documents using XPath expressions with built-in caching for performance.

#caching#xpath-selector#html-parsing
Stars784
Forks80
Last commit22 days ago
image-scraper
image-scraperPython

A high-performance, multithreaded command-line tool for downloading images from webpages.

#pypi#commandline-tool#terminal
Stars776
Forks104
Last commit8 years ago
Essence
EssencePHP

A PHP library for extracting media information from web pages like YouTube videos, Twitter statuses, and blog articles.

#metadata-parsing#url-crawling#php-library
Stars772
Forks80
Last commit3 years ago
xpath
xpathGo

A Go package for querying XML, HTML, and JSON documents using XPath expressions.

#xpath-query#selects-descendants#document-query
Stars743
Forks98
Last commit1 day ago
fattest-cat
fattest-catJavaScript

A script to find the fattest cat currently available for adoption at the San Francisco SPCA.

#fun-project#san-francisco#pet-adoption
Stars736
Forks38
Last commit2 years ago
dataflowkit
dataflowkitGo

A Go web scraping framework that extracts structured data from websites using CSS selectors, including JavaScript-rendered pages.

#chrome-fetcher#scraping-websites#javascript-rendering
Stars715
Forks83
Last commit3 years ago
Hickory
HickoryClojure

A Clojure/ClojureScript library that parses HTML into Clojure data structures for analysis, transformation, and serialization.

#dom-manipulation#hiccup#clojurescript
Stars678
Forks55
Last commit3 months ago
pychrome
pychromePython

A Python package for controlling Google Chrome/Chromium via the Chrome DevTools Protocol with a threading-based API.

#threading#chrome#headless-chrome
Stars647
Forks116
Last commit2 years ago
tor-browser-selenium
tor-browser-seleniumPython

A Python library for automating Tor Browser with Selenium WebDriver for privacy-focused web scraping and testing.

#selenium#privacy#onion-services
Stars594
Forks104
Last commit1 year ago
Steam Community
Steam CommunityJavaScript

A Node.js library for interacting with Steam Community's website interfaces, including login, trading, and inventory management.

#trading-bot#steam#steam-community
Stars575
Forks156
Last commit11 days ago
LinkThumbnailer
LinkThumbnailerRuby

Ruby gem that fetches images and metadata from URLs to generate link previews, similar to social media previews.

#content-parsing#thumbnail-generation#metadata-extraction
Stars510
Forks105
Last commit1 year ago
playwright-ruby-client
playwright-ruby-clientRuby

A Ruby client library for browser automation and testing using Microsoft Playwright.

#playwright#ruby-gem#headless-browser
Stars499
Forks52
Last commit3 days ago
HLTV
HLTVTypeScript

An unofficial Node.js API for programmatically accessing HLTV's Counter-Strike esports data, including matches, teams, players, and live scores.

#statistics#live-scores#scraper
Stars493
Forks126
Last commit1 year ago
azuretls-client
azuretls-clientGo

A Go HTTP client that spoofs TLS/JA3, HTTP/2, and HTTP/3 fingerprints to emulate real browsers by default.

#proxy-support#ja3-fingerprint#browser-emulation
Stars464
Forks65
Last commit3 months ago
deno-puppeteer
deno-puppeteerTypeScript

A port of the Puppeteer browser automation library to run natively on Deno.

#puppeteer#screenshot#headless-chrome
Stars458
Forks47
Last commit2 years ago
Android-Link-Preview
Android-Link-PreviewJava

An Android library that generates link previews by extracting titles, descriptions, and images from URLs.

#content-preview#metadata-extraction#link-preview
Stars414
Forks130
Last commit6 years ago
Nokolexbor
NokolexborC

A high-performance, Nokogiri-compatible HTML5 parser for Ruby with CSS selector and XPath support.

#dom-manipulation#css-selectors#html5
Stars414
Forks8
Last commit22 days ago
Lambda Soup
Lambda SoupOCaml

A functional HTML scraping and manipulation library for OCaml with CSS selector support.

#ocaml-library#functional-programming#css-selectors
Stars409
Forks35
Last commit1 year ago
godet
godetGo

A Go client library for remotely controlling Chrome/Chromium browsers via the Chrome DevTools Protocol.

#go-client#go-library#headless-browser
Stars398
Forks43
Last commit4 months ago
Web Scraping Reference: Cheat Sheet for Web Scraping using R
Web Scraping Reference: Cheat Sheet for Web Scraping using RR

A comprehensive cheat sheet and reference for web scraping in R using rvest, httr, and RSelenium.

#r-programming#webscraping#httr
Stars397
Forks101
Last commit
Gorilla
GorillaRust

A versatile Rust tool for generating and mutating wordlists using patterns, web scraping, and password formats.

#cracking#hash#infosec
Stars390
Forks22
Last commit5 months ago
hget
hgetHTML

A CLI and API tool that converts HTML into plain text, Markdown, or filtered HTML for terminal viewing.

#developer-tools#terminal-utility#content-extraction
Stars388
Forks13
Last commit2 years ago
rookie
rookieRust

A cross-platform library to load and decrypt cookies from any web browser, built with Rust for speed and safety.

#cookies#privacy-tools#cookie-extraction
Stars364
Forks48
Last commit6 months ago
read-art
read-artJavaScript

A Node.js library to automatically scrape and extract readable article content from any web page, supporting both English and Chinese.

#readability#content-extraction#crawler
Stars346
Forks36
Last commit8 years ago
crawley
crawleyGo

A fast, Unix-style command-line web crawler that extracts links, resources, and API endpoints from web pages.

#api-discovery#resource-discovery#link-extraction
Stars340
Forks18
Last commit7 days ago
scrape
scrapeElixir

An Elixir library for structured data extraction from websites, articles, and RSS/Atom feeds using information-retrieval techniques.

#readability#elixir#information-retrieval
Stars337
Forks41
Last commit6 years ago
meseeks
meseeksElixir

An Elixir library for parsing and extracting data from HTML and XML using CSS or XPath selectors.

#elixir#css-selectors#html5
Stars325
Forks26
Last commit1 year ago
meeseeks
meeseeksElixir

An Elixir library for parsing and extracting data from HTML and XML using CSS or XPath selectors.

#elixir#css-selectors#nif
Stars325
Forks26
Last commit1 year ago
wikipedia
wikipediaRuby

A Ruby client library for interacting with the Wikipedia API, providing easy access to articles, summaries, images, and metadata.

#content-parsing#data-fetching#ruby-gem
Stars309
Forks74
Last commit3 years ago
PreviousPage 3 of 6

Related Tags

Community-curated · Updated weekly · 100% open source

Found a gem we're missing?

Open-Awesome is built by the community, for the community. Submit a project, suggest an awesome list, or help improve the catalog on GitHub.

Submit a projectStar on GitHub
3 years ago
Next
#Automation47
#Browser Automation45
#Testing38
#Data Extraction35
#Headless Browser32
#Crawler31
#Http Client24
#Go23
#Python22
#Headless Chrome22
#Chrome21
#Golang20