Showing 32 of 32 projects
A Python utility for converting PDFs, Office documents, images, audio, and more into structured Markdown for LLM consumption.
A collection of utilities that help customize Windows and streamline everyday tasks.
A command-line tool that adds an OCR text layer to scanned PDF files, making them searchable and copy-pasteable.
A pure-Python PDF library for splitting, merging, cropping, transforming, and extracting data from PDF files.
A line-oriented search tool that extends ripgrep to search inside PDFs, Office documents, archives, and many other file types.
A polyglot document intelligence framework with a Rust core for extracting text, metadata, and structured data from 91+ file formats.
A Python library for extracting and analyzing text, images, and metadata from PDF documents.
A Python library and CLI tool for web crawling, scraping, and extracting main text, metadata, and comments from web pages.
A Python library and CLI tool for automatic text summarization using extractive methods like LexRank, LSA, Luhn, and Edmundson.
A pure JavaScript OCR engine compiled from Ocrad via Emscripten for client-side text recognition in the browser.
A Go package for Optical Character Recognition (OCR) using the Tesseract C++ library.
A curated list of awesome open-source OCR software, libraries, datasets, and literature.
A macOS menu bar app that uses OCR to copy any text visible on your screen directly to your clipboard.
A Java JNA wrapper for Tesseract OCR API, enabling OCR functionality in Java applications.
A lightweight Linux desktop application that extracts text from images using OCR with drag-and-drop simplicity.
A comprehensive natural language processing framework for Ruby with support for text extraction, parsing, and machine learning.
A high-performance PDF toolkit for text/image extraction, markdown conversion, and PDF editing, built in Rust with Python, WASM, CLI, and MCP server bindings.
A simple OCR API server that's easy to deploy with Docker or on Heroku.
A Ruby library for extracting text and metadata from various document formats using Apache Tika.
A .NET framework for extracting and exporting text and data from a wide variety of document formats.
A Ruby gem for extracting pages from PDFs as images and text strings using Ghostscript, ImageMagick, and pdftotext.
Archived R package for accessing the Monkeylearn API for text classification and extraction.
A high-performance Go library for PDF creation, reading, table extraction, digital signatures, and encryption.
Enables full-text search within uploaded documents (PDF, Word, Excel) in Wagtail CMS.
A V programming language wrapper for Tesseract-OCR, enabling text extraction and OCR operations from images.
Extracts hyperlinks from files using Apache Tika for batch processing and web archiving workflows.
A Blazor Server app that uses Azure Computer Vision to extract printed text from uploaded images.
A Crystal library that extracts text and formatting from .DOCX files and converts them to Markdown.
A Python parser for HandyGames .lng files, converting them to JSON for translation.
A TypeScript-first, framework-agnostic toolkit for scanning, extracting, and managing i18n translations with CSV support.
A WezTerm plugin that extracts text from command output and inserts it into the next prompt.
Angular web app that generates alt texts and transcribes text in images using AI models from OpenAI and Google.
Open-Awesome is built by the community, for the community. Submit a project, suggest an awesome list, or help improve the catalog on GitHub.