Showing 3 of 3 projects
An unsupervised text tokenizer and detokenizer for neural network-based text generation systems with subword units.
Fast, state-of-the-art tokenizers for training and tokenization, optimized for both research and production.
A Rust implementation of OpenAI's tiktoken tokenizer for working with GPT models and token counting.
Open-Awesome is built by the community, for the community. Submit a project, suggest an awesome list, or help improve the catalog on GitHub.