InferrLM - On-device AI for iOS & Android
-
Updated
Jun 29, 2026 - TypeScript
InferrLM - On-device AI for iOS & Android
Provides an offline AI tutoring application built with Electron that runs locally on your computer using Ollama or Llama models, ensuring complete privacy without requiring internet connectivity.
Local-first AI workspace for macOS and Windows with cloud providers, local inference, MCP tools, memory, and artifacts
An LLM interface for your Ollama models, running entirely offline. No cloud, no subscription or API costs. Experience local AI on your own terms.
An opinionated distribution of pi, the coding agent. Extensions for project memory, spec-driven development, local LLM inference, parallel task decomposition, and more.
A minimal, local-first LLM harness extracted from opencode's infrastructure
A 90s drug empire simulation powered by local LLM inference. 96 AI dealers, 8 drugs, 9 cities, 5 tiers, cops, cartels, prison, girlfriends, crew — all running on your GPU via llama.cpp.
Local AI inference, embedded in your app. One OpenAI-compatible API across Web, Node, iOS, Android, React Native, Flutter, .NET, and Capacitor. Ten backends (llama.cpp · MediaPipe · MLX · Transformers.js · WebLLM · LiteRT · LiteRT-LM · ONNX · ML.NET · llamafile). Distributed inference across the devices your users own.
Workplane is a control plane for shell tasks, local inference, and AI agent harnesses — on your laptop, home server, and GPU box. Compose multi-step workplans that mix local Ollama with frontier APIs, over Tailscale, WireGuard, or a private LAN.
Jeopardy! emulator with API support for local (Ollama) or external (OpenRouter) inference (Ollama) to batch-generate categories and clues of a board. Includes a system prompt and user input field
Lets browser tabs share one on-device AI model safely, with queues, streaming, cancellation, takeover and runtime integrity checks.
A zero-cost, parasitic AI engine that intercepts LLM API calls and executes them locally on the client's hardware using WebAssembly and WebGPU.
ElizaOS v1.x agent running Gemma 3 27B locally via Ollama on an RTX 3090, dogfooding @thecolony/elizaos-plugin against The Colony (thecolony.cc).
agent for groupchats, optimized for tact/silence and privacy + specially finetuned local model(s) for it
AI-assisted constituency case writer — causality engine, multi-agency letter generation, and human-in-the-loop governance. Runs on local inference.
A full-scale Golden Age of Piracy AI agent simulation powered by local LLM inference via llama.cpp. 100+ autonomous AI agents, 15 agent types, 50+ trade goods, 25 sea zones, and rich narrative proprioception — all running on your GPU.
Add a description, image, and links to the local-inference topic page so that developers can more easily learn about it.
To associate your repository with the local-inference topic, visit your repo's landing page and select "manage topics."