Distribute and run LLMs with a single file.
-
Updated
Jul 27, 2026 - C++
Distribute and run LLMs with a single file.
High-speed Large Language Model Serving for Local Deployment
Mano-P: Open-source GUI-VLA agent for edge devices. #1 on OSWorld (specialized, 58.2%). Runs locally on Apple M4 Mac mini/MacBook — no data leaves your device.Mano-P 是一个开源 GUI-VLA 项目,支持在 Mac mini/MacBook 上或通过算力棒本地运行推理,实现纯视觉驱动的跨平台 GUI 自动化操作。数据完全本地处理,支持复杂多步骤任务规划与执行。
[ICLR'25] Fast Inference of MoE Models with CPU-GPU Orchestration
Modern desktop application (Rust + Tauri v2 + Svelte 5 + Candle (HF)) for communicating with AI models that runs completely locally on your computer. No subscriptions, no data sent to the internet — just you and your personal AI assistant
InferrLM - On-device AI for iOS & Android
Notolog Markdown Editor
A fully browser-native RAG application for document Q&A, powered by Rust and WebAssembly with local vector search, embeddings, and in-browser LLM inference.
A lightweight CUDA-based local inference platform built around Z-Image Turbo by Tongyi
Local AI music generator with smart lyrics: Gradio web UI for HeartMuLa + Ollama/OpenAI, tags, history, and high-fidelity audio.
An open-source, model-agnostic agent harness for local LLMs. Define agents in YAML (tools, memory, deny-first permissions) and run them against any OpenAI-compatible endpoint: vLLM, Ollama, LM Studio, or llama.cpp.
Tool for test diferents large language models without code.
Provides an offline AI tutoring application built with Electron that runs locally on your computer using Ollama or Llama models, ensuring complete privacy without requiring internet connectivity.
Honeycomb Lab — hex map + OpenAI gateway control plane for a home AI fleet
Local inference server for Apple Silicon — hot-swaps MLX models (LLM, vision, embeddings, TTS, STT) via OpenAI API
.NET library for on-demand local AI model inference — zero bundled models, lazy loading, hardware-aware GPU/CPU selection, 10 task types including embeddings, generation, vision, and audio.
Edge Agent Lab is an Android testing platform for evaluating small language model (SLM) agents directly on mobile devices.
local inference, fully under your hand.
LLM chatbot example using OpenVINO with RAG (Retrieval Augmented Generation).
Local LLM inference on Xbox Series S|X (UWP) via ONNX Runtime GenAI.
Add a description, image, and links to the local-inference topic page so that developers can more easily learn about it.
To associate your repository with the local-inference topic, visit your repo's landing page and select "manage topics."