Showing 23 of 23 projects
Official JAX/Flax implementation of Vision Transformer (ViT) and MLP-Mixer for image recognition, with pre-trained models.
A cross-platform Python module for programmatically controlling the mouse and keyboard to automate GUI interactions.
Fast and simple OCR library for iOS/macOS using neural networks, optimized for short alphanumeric codes.
Open source robotic process automation software for automating repetitive tasks across desktop and web applications.
A CNN-based captcha solver for Taiwan Railway booking website with a training set generator that mimics captcha style and uses data augmentation.
A TensorFlow-based image recognition system for captchas that works without image segmentation.
Android app that uses your camera to identify objects and translate their names into different languages.
A deep learning framework for detecting and localizing upper-body, lower-body, and full-body clothes in fashion images.
A Flutter plugin for integrating Firebase ML Kit's on-device and cloud machine learning features into mobile apps.
A convolutional neural network for CAPTCHA recognition using Keras and PyTorch.
Course materials for GWU's Data Mining and Machine Learning classes covering preprocessing, modeling, and practical Kaggle applications.
A Python-based CAPTCHA breaking solution using Keras and OpenCV, developed for a data science competition.
A YOLO-based object detection system specifically trained to identify DJI drones in images and video.
A fast Elixir library to parse image binaries and extract dimensions, mime-type, and validity for 13+ formats.
A deprecated Go client library for the Clarifai v1 API, enabling image tagging and feedback functionality.
Extends Selenium Grid with Sikuli image recognition, file upload/download, and remote UI automation capabilities.
TensorFlow implementation of AlexNet with 3D convolutional layers for volumetric image recognition.
An Adobe Air Native Extension for building augmented reality apps on Android and iOS using the Wikitude SDK.
A legacy Ruby SDK for accessing AlchemyAPI's text analysis and image recognition AI services.
A WebDriver-compatible server that provides Selenium WebDriver API access to AutoPy for cross-platform image recognition and UI automation.
A Blazor Server app that uses Azure Computer Vision to extract printed text from uploaded images.
A complete neural network solution for handwritten digit recognition, built with Ruby and featuring a web interface.
A decentralized CNN classification service that distributes pre-trained image models across Golem network nodes for scalable predictions.
Open-Awesome is built by the community, for the community. Submit a project, suggest an awesome list, or help improve the catalog on GitHub.