Showing 36 of 74 projects
An offline, open-source voice assistant toolkit for home automation, supporting multiple languages and privacy-focused self-hosting.
:speech_balloon: SpeechPy - A Library for Speech Processing and Recognition: http://speechpy.readthedocs.io/en/latest/
Swift SDK for integrating IBM Watson AI services like speech, language, and assistant into iOS and Linux applications.
An Emacs org-mode minor mode that integrates generative AI models like ChatGPT, DALL-E, and Stable Diffusion for text and image generation.
A Node.js library for adding voice interfaces with offline hotword detection and cloud speech recognition.
An Android Input Method Editor (IME) providing offline voice recognition and translation using the Whisper engine.
A Unity SDK for integrating IBM Watson AI services like speech, language, and vision into games and applications.
A customizable iOS overlay that handles voice permission and converts speech to text using native speech recognition.
A joint audio tagging and speech recognition model that adds audio event detection to OpenAI Whisper with minimal computational overhead.
A Google Colab notebook that transcribes YouTube videos using OpenAI's Whisper speech recognition model.
Control Vim and Vim-like editors with voice commands using speech recognition.
A command-line interface for blazingly fast audio transcription using optimized Whisper ASR models.
An open-source desktop app that transcribes voice to polished text using AI and types it into any application.
A curated collection of linguistic resources, tools, and datasets for Natural Language Processing and Computational Linguistics on Spanish.
A curated collection of linguistic resources, datasets, and tools for Natural Language Processing and Computational Linguistics on Spanish.
An on-device AI teleprompter that listens to your conversations and suggests charismatic quotes in real-time.
A Flutter plugin for speech recognition on iOS and Android using native APIs.
A curated collection of datasets, corpora, and resources for Indonesian natural language processing tasks.
An Android overlay that handles voice permission and converts user speech to text with a customizable UI.
Ruby FFI bindings for Pocketsphinx, a lightweight speech recognition engine.
A Chrome extension that adds hands-free voice control to ChatGPT with custom trigger phrases and 60+ language support.
An Android chatbot with voice interaction capabilities powered by IBM Watson's AI services on IBM Cloud.
A fork of OpenAI's Whisper speech recognition models optimized with OpenVINO backend for faster CPU inference.
A collection of iOS sample apps demonstrating Generative AI capabilities including OpenAI, local LLMs, Stable Diffusion, and speech recognition.
Android client library for integrating IBM Watson cognitive services like speech recognition, text-to-speech, and visual recognition.
A Python library for easy access, management, and processing of audio datasets, particularly for machine learning tasks.
A Capacitor plugin for native speech recognition on iOS and Android, enabling voice-to-text in hybrid mobile apps.
A collection of refactored, high-quality Android examples demonstrating TensorFlow Lite for on-device machine learning tasks.
A deep learning system for automatic spoken language identification from audio files using TensorFlow and Caffe.
A Python library for unsupervised learning of hidden semi-Markov models with explicit durations.
A Docker-based speech recognition model that converts short English WAV audio files into text using Mozilla's DeepSpeech.
A speech-to-text module for Godot 3 that captures microphone input and converts it to text for game development.
A Ruby library for consuming the AT&T Speech API to convert speech to text and text to speech.
A Capacitor plugin providing natural, low-latency speech recognition for iOS and Android apps with streaming results and permission helpers.
An unofficial Elixir SDK for Microsoft Azure Speech Service, providing speech-to-text and text-to-speech capabilities.
An Angular directive that provides an easy-to-use wrapper for the Web Speech API for voice input.
Open-Awesome is built by the community, for the community. Submit a project, suggest an awesome list, or help improve the catalog on GitHub.