Showing 8 of 8 projects
An unsupervised text tokenizer and detokenizer for neural network-based text generation systems with subword units.
Fast, state-of-the-art tokenizers for training and tokenization, optimized for both research and production.
A Python library and CLI tool for web crawling, scraping, and extracting main text, metadata, and comments from web pages.
A Python NLP library built on spaCy for text preprocessing, feature extraction, and analysis tasks.
A Julia package providing high-performance, configurable tokenizers and sentence splitters for natural language processing.
A rule-based Unicode tokenizer that separates words from punctuation and splits sentences for NLP preprocessing.
A scikit-learn pipeline implementing the projection layer of Self-Governing Neural Networks (SGNN) using character n-grams and random hashing.
A Python package for cleaning text data for NLP tasks, including language filtering, duplicate removal, and outlier detection.
Open-Awesome is built by the community, for the community. Submit a project, suggest an awesome list, or help improve the catalog on GitHub.