Showing 35 of 35 projects
A curated list of resources, tools, datasets, and learning materials for Chinese Natural Language Processing.
A multi-domain Chinese word segmentation toolkit offering higher accuracy and domain-specific models.
A curated list of 100 foundational and influential papers in natural language processing for students and researchers.
A Python NLP library built on spaCy for text preprocessing, feature extraction, and analysis tasks.
A Python library and CLI tool for converting text to phonetic transcriptions (phones) across multiple languages using various backends.
A comprehensive natural language processing framework for Ruby with support for text extraction, parsing, and machine learning.
A curated list of awesome resources, libraries, and tools for natural language processing (NLP) in Ruby.
A curated list of awesome resources, libraries, and tools for natural language processing (NLP) in Ruby.
A curated directory of academic institutions and principal investigators in computational neuroscience worldwide.
An R package for the quantitative analysis of textual data, providing comprehensive tools for natural language processing and text management.
A curated list of open-access resources and tools for Natural Language Processing (NLP) focused on the German language.
An AI system that incrementally generates scientific paper drafts by predicting links between concepts and generating text sections.
PyNLPl, pronounced as 'pineapple', is a Python library for Natural Language Processing. It contains various modules useful for common, and less common, NLP tasks. PyNLPl can be used for basic tasks such as the extraction of n-grams and frequency lists, and to build simple language model. There are also more complex data types and algorithms. Moreover, there are parsers for file formats common in NLP (e.g. FoLiA/Giza/Moses/ARPA/Timbl/CQL). There are also clients to interface with various NLP specific servers. PyNLPl most notably features a very extensive library for working with FoLiA XML (Format for Linguistic Annotation).
A high-performance Go library for calculating Levenshtein distance between strings, including Unicode support.
A curated list of resources, tools, datasets, and communities for linguistics and natural language processing.
A curated collection of linguistic resources, tools, and datasets for Natural Language Processing and Computational Linguistics on Spanish.
A curated collection of linguistic resources, datasets, and tools for Natural Language Processing and Computational Linguistics on Spanish.
A curated list of free tools, datasets, models, and resources for Hungarian Natural Language Processing.
BLLIP reranking parser (also known as Charniak-Johnson parser, Charniak parser, Brown reranking parser) See http://pypi.python.org/pypi/bllipparser/ for Python module.
A Java library for parsing and generating text using combinatory categorial grammar and hybrid logic dependency semantics.
A statistical natural language generator for spoken dialogue systems, supporting both A*-search and seq2seq algorithms.
A C++ and Python library for efficient extraction and analysis of n-grams, skipgrams, and flexgrams from large corpora.
A natural language processing library for Uralic and other languages, offering morphological analysis, generation, lemmatization, and lexical information.
A Julia package providing high-performance, configurable tokenizers and sentence splitters for natural language processing.
A tagger, lemmatizer, morphological analyzer, and dependency parser for Dutch using memory-based NLP modules.
A rule-based Unicode tokenizer that separates words from punctuation and splits sentences for NLP preprocessing.
A CCG parser implementing all combinators with parsing to logical form and parameter estimation for probabilistic CCG.
This is a Python binding to the tokenizer Ucto. Tokenisation is one of the first step in almost any Natural Language Processing task, yet it is not always as trivial a task as it appears to be. This binding makes the power of the ucto tokeniser available to Python. Ucto itself is regular-expression based, extensible, and advanced tokeniser written in C++ (http://ilk.uvt.nl/ucto).
A Common Lisp implementation of Combinatory Categorial Grammar (CCG) with a full set of combinators and probabilistic extensions.
A Java library for bilingual English/French text surface realization, adapted from SimpleNLG v4.2.
A Java API for generating German natural language text, adapted from SimpleNLG 4.
Ruby based API for the project Wortschatz Leipzig.
A C++ library for reading, manipulating, and creating FoLiA (Format for Linguistic Annotation) documents.
A collection of Latin American corpora, dictionaries, and text resources for natural language processing and text mining.
A web-based annotation platform for Combinatory Categorial Grammar (CCG) parsing and treebank creation.
Open-Awesome is built by the community, for the community. Submit a project, suggest an awesome list, or help improve the catalog on GitHub.