Showing 7 of 7 projects
A C/C++ library for efficient, cross-platform LLM inference with extensive hardware support and quantization.
A lightweight, single-binary Rust inference server providing 100% OpenAI-API compatible endpoints for local GGUF models.
An open-source ChatGPT alternative that runs 100% offline on your computer with local LLMs or cloud model connections.
A library for running LLMs locally and efficiently on any device with support for Python, Flutter, and Godot.
Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.
A native desktop application for running large language models locally with Rust and Tauri, offering private AI chat without internet or cloud services.
Elixir NIF wrapper for llama_cpp.rs enabling GGUF model inference in Elixir/Erlang applications.
Open-Awesome is built by the community, for the community. Submit a project, suggest an awesome list, or help improve the catalog on GitHub.