Showing 36 of 392 projects
A TensorFlow-based model that generates text summaries using pointer-generator networks, deployable as a Docker container.
A Rust crate for generating n-grams from token sequences with Unicode padding support.
A Common Lisp implementation of Combinatory Categorial Grammar (CCG) with a full set of combinators and probabilistic extensions.
An Elixir port of Nakatani Shuyo's natural language detection library, supporting 55 languages.
A Haxe library for linguistical analysis and natural language processing with tokenization, stemming, classification, and dictionary features.
A Go client library for the Detect Language API, enabling language detection and account management.
A TensorFlow-based model that generates English-language text similar to news articles from the One Billion Word Benchmark dataset.
NALP is a Python library for natural language processing and adversarial learning, from embeddings to neural networks.
A named entity recognition model that locates and tags entities like persons, locations, and organizations in text using a neural network.
A repository for planning and training German transformer language models from scratch.
Elixir implementation of Simhash for text similarity detection using N-gram features.
A scikit-learn pipeline implementing the projection layer of Self-Governing Neural Networks (SGNN) using character n-grams and random hashing.
Interactive lecture notes on probabilistic topic models using Jupyter notebooks, covering LDA, Dirichlet processes, and inference methods.
Go bindings for the Snowball libstemmer library, providing stemming algorithms like Porter and Porter2.
A Go tokenizer for Chinese text segmentation using dictionary and Bigram language models.
Natural language processing algorithms implemented in pure Ruby with minimal dependencies.
A Crystal library for building Markov Chains and running Markov Processes.
German language versions of GPT-2, trained on the CC-100 corpus and initialized from English GPT-2 weights.
A multilingual Part-of-Speech tagger and lemmatizer for Basque, Dutch, English, French, Galician, German, Italian, and Spanish.
Ruby based API for the project Wortschatz Leipzig.
A Ruby gem for performing sentiment analysis on German text using dictionary-based scoring.
An Angular component library for interactively highlighting and annotating text, ideal for visualizing named entity recognition or part-of-speech tagging.
Python binding for Morfologik, a Polish morphological analyzer and stemmer.
A C++ library for reading, manipulating, and creating FoLiA (Format for Linguistic Annotation) documents.
An Elixir library for calculating tf-idf (term frequency–inverse document frequency) scores to identify important words in text.
A Neovim plugin that uses LLMs to generate and execute Vim commands from natural language prompts.
A Python command-line tool for encoding text sequences into vector-of-vector representations using word embeddings.
A Go implementation of the Porter stemming algorithm for English text, offering both simple and zero-allocation APIs.
A Java-based machine learning framework for building and training classifiers with a declarative syntax.
A simple and extensible Ruby gem for sentiment analysis with customizable analysis strategies.
An Elixir library for translating between Japanese scripts (hiragana, katakana, romaji, kanji) and performing morphological analysis using Mecab.
A research project exploring multilingual BERT models for Named Entity Recognition in German and English using the CoNLL-2003 dataset.
Add natural language search to Wagtail images using OpenAI's CLIP model.
Go binding for the libtextcat C library, providing language detection capabilities.
A Swift playground for sentiment analysis using AFINN-165 wordlist and Emoji Sentiment Ranking.
An Angular directive that converts natural language to dates using Chrome's built-in AI Writer API.
Open-Awesome is built by the community, for the community. Submit a project, suggest an awesome list, or help improve the catalog on GitHub.