Showing 21 of 21 projects
An unsupervised text tokenizer and detokenizer for neural network-based text generation systems with subword units.
A minimalistic, single-header JSON tokenizer/parser in C for resource-limited and embedded systems.
A blazing fast and feature-rich parser building toolkit for JavaScript, supporting LL(K) and LL(*) grammars.
A high-performance, browser-grade HTML5 parser written in Rust, developed as part of the Servo project.
A self-contained Japanese morphological analyzer written in pure Go, tokenizing text into words and analyzing parts of speech.
A Swift library for tokenizing strings using character sets and custom tokenizers when whitespace splitting is insufficient.
A comprehensive suite of Java NLP libraries and tools for text annotation, feature extraction, and language processing tasks.
A multilingual command-line sentence tokenizer written in Go, ported from NLTK's Punkt system.
A Rust implementation of OpenAI's tiktoken tokenizer for working with GPT models and token counting.
Bayesian text classifier for Go with flexible tokenizers and storage backends.
A high-performance, regex-free Go tokenizer for parsing strings, slices, and infinite streams into customizable tokens.
A simple tokenizer in Ruby for NLP tasks.
This is a Python binding to the tokenizer Ucto. Tokenisation is one of the first step in almost any Natural Language Processing task, yet it is not always as trivial a task as it appears to be. This binding makes the power of the ucto tokeniser available to Python. Ucto itself is regular-expression based, extensible, and advanced tokeniser written in C++ (http://ilk.uvt.nl/ucto).
A Ruby implementation of Naive Bayes text classification designed as a strategy for the OmniCat classification framework.
A native YAML parser for the V programming language, supporting reading, tokenization, and conversion to JSON or dynamic structures.
A multi-language tokenizer for extracting identifiers from source code using tree-sitter and pygments parsers.
A Go tokenizer for Chinese text segmentation using dictionary and Bigram language models.
A simple tokenizer library for Deno that breaks text into tokens using customizable rules.
A Go package for filtering specific words from text using customizable tokenizers and filters.
Crystal implementation of the TOON format, a compact serialization format designed to reduce token usage for LLM input.
A command-line tool for inspecting V source files by printing AST and tokens.
Open-Awesome is built by the community, for the community. Submit a project, suggest an awesome list, or help improve the catalog on GitHub.