A wannabe Ollama equivalent for Apple MlX models
-
Updated
Mar 2, 2025 - Shell
A wannabe Ollama equivalent for Apple MlX models
🚀 The Ultimate Curated List of LLMOps Tools, Frameworks, and Resources - A comprehensive collection of the best tools for Large Language Model Operations
Docker compose configs for serving Qwen3.8-27B on a DGX Spark (GB10): vLLM+MTP current, SGLang+DSPARK + FP8 rollback stacks
Production-ready checklists and frameworks for deploying LLMs, GenAI models, and AI infrastructure. Covers vLLM, Kubernetes, GPU optimization, observability, compliance, and Day-0 to Day-2 operations.
Run GPT-OSS 120B on NVIDIA DGX Spark with vLLM, build an API server, and create a local AI coding assistant
Enterprise-grade, high-performance LLM serving on Kubernetes. OpenAI-compatible inference powered by vLLM and OpenVINO™ Model Server behind a unified AI gateway, with NUMA-aware CPU optimization for Intel® Xeon® - production model serving, routing, and security built in.
Production-ready platform for deploying OpenAI-compatible LLM inference on Kubernetes. A one-command installer for enterprise generative AI on Intel® Xeon®, with vLLM and other serving engines, secure Keycloak authentication, and built-in observability - from bare metal to a scalable AI endpoint.
Enterprise-grade, high-performance LLM serving on Kubernetes. OpenAI-compatible inference powered by vLLM and OpenVINO™ Model Server behind a unified AI gateway, with NUMA-aware CPU optimization for Intel® Xeon® — production model serving, routing, and security built in.
Customize Nvidia Triton to use OpenShift Source to Image building
End-to-end MLOps platform: Terraform + K3s + ArgoCD + MLflow + BentoML + Evidently AI + Prometheus/Grafana — self-hosted on a homelab
Production-ready platform for deploying OpenAI-compatible LLM inference on Kubernetes. A one-command installer for enterprise generative AI on Intel® Xeon®, with vLLM and other serving engines, secure Keycloak authentication, and built-in observability — from bare metal to a scalable AI endpoint.
Local LLM inference lab with llama.cpp-compatible serving, GGUF model profiles, OpenAI-compatible API checks, and OpenWebUI integration.
Achieve state of the art inference performance with modern accelerators on Kubernetes
Add a description, image, and links to the model-serving topic page so that developers can more easily learn about it.
To associate your repository with the model-serving topic, visit your repo's landing page and select "manage topics."