Showing 10 of 10 projects
Run large language models (LLMs) privately on everyday desktops and laptops without requiring API calls or GPUs.
An open-source, locally-runnable code completion engine using large language models that works on CPU.
A C#/.NET library for efficient local inference of LLaMA and other large language models, based on llama.cpp.
A library for running LLMs locally and efficiently on any device with support for Python, Flutter, and Godot.
A terminal user interface (TUI) for interacting with multiple large language model backends, featuring Vim keybindings and chat history.
A minimalistic C++ Jinja templating engine specifically designed for LLM chat templates, used in llama.cpp and other projects.
Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.
A one-command installer that transforms a fresh Arch Linux installation into a complete, local-first AI development environment with Hyprland.
Elixir NIF wrapper for llama_cpp.rs enabling GGUF model inference in Elixir/Erlang applications.
A local compatibility server that provides a subset of the Gemini and OpenAI APIs for testing and development, backed by a local LLM.
Open-Awesome is built by the community, for the community. Submit a project, suggest an awesome list, or help improve the catalog on GitHub.