Showing 36 of 225 projects
A BERT-based model that detects six types of toxicity in text comments, deployable as a Docker container.
A text stream monitoring library that detects keywords, phrases, regexes, and complex Lucene queries in documents.
🕶️ Adds chat auto-clear functionality to ChatGPT for more privacy
A Python library for blazing fast fuzzy and semantic text search using vectorization and approximate nearest neighbor techniques.
A high-performance Porter2 stemmer implementation using finite state machines for suffix comparison.
A simple tokenizer in Ruby for NLP tasks.
A collection of samples demonstrating advanced usage of context variables and system entities in IBM Watson Assistant.
A phonetic algorithm for indexing Chinese characters by sound, estimating distance between words and finding similar-sounding candidates.
A Ruby library implementing the Earley parsing algorithm for any context-free language, with support for ambiguous grammars and parse forests.
A Ruby gem for calculating text readability statistics, complexity metrics, and grade levels across 22 languages with high performance.
CGo bindings for Yandex.Mystem, providing Russian morphological analysis in Go applications.
A multilingual Rust implementation of the RAKE algorithm for automatic keyword extraction from text.
A Julia package providing lazy-loading iterators for various NLP corpora with automatic data dependency management.
This is a Python binding to the tokenizer Ucto. Tokenisation is one of the first step in almost any Natural Language Processing task, yet it is not always as trivial a task as it appears to be. This binding makes the power of the ucto tokeniser available to Python. Ucto itself is regular-expression based, extensible, and advanced tokeniser written in C++ (http://ilk.uvt.nl/ucto).
A multilingual library for parsing natural language date strings into java.util.Date objects.
A Ruby gem for customizable text tokenization, useful for web crawling and natural language processing.
A German ELMo deep contextualized word representation model trained on a specialized German Wikipedia text corpus for NLP tasks.
A JavaScript bilingual text realizer for web development, generating French and English text with HTML integration.
A Haxe library for linguistical analysis and natural language processing with tokenization, stemming, classification, and dictionary features.
A curated collection of transformers to accelerate and enhance machine learning experimentation with the Steppy library.
A curated list of the best production-ready Hugging Face models for NLP, vision, audio, and multimodal tasks.
German language versions of GPT-2, trained on the CC-100 corpus and initialized from English GPT-2 weights.
Natural language processing algorithms implemented in pure Ruby with minimal dependencies.
A Java API for generating German natural language text, adapted from SimpleNLG 4.
Ruby based API for the project Wortschatz Leipzig.
A Ruby gem for performing sentiment analysis on German text using dictionary-based scoring.
Python binding for Morfologik, a Polish morphological analyzer and stemmer.
A C++ library for reading, manipulating, and creating FoLiA (Format for Linguistic Annotation) documents.
A GitHub Action that monitors PR/issue comments and warns users who use offensive language.
An Angular component library for interactively highlighting and annotating text, ideal for visualizing named entity recognition or part-of-speech tagging.
A Python package for cleaning text data for NLP tasks, including language filtering, duplicate removal, and outlier detection.
A Rust crate providing embeddings and positional encoding implementations for NLP and Transformer-based models.
Generate word embedding vectors from text files using the Swivel algorithm on IBM Watson Machine Learning.
A Ruby framework for Hebrew-English transliteration using customizable phoneme maps.
Parses intent utterance files like the Alexa Skills Kit Sample Utterance format to extract intents, slots, and words.
A Go microservice that provides sentiment analysis for text using the VADER algorithm.
Open-Awesome is built by the community, for the community. Submit a project, suggest an awesome list, or help improve the catalog on GitHub.