Skip to content
#

model-serving

Here are 13 public repositories matching this topic...

Production-ready checklists and frameworks for deploying LLMs, GenAI models, and AI infrastructure. Covers vLLM, Kubernetes, GPU optimization, observability, compliance, and Day-0 to Day-2 operations.

  • Updated Aug 25, 2026
  • Shell

Enterprise-grade, high-performance LLM serving on Kubernetes. OpenAI-compatible inference powered by vLLM and OpenVINO™ Model Server behind a unified AI gateway, with NUMA-aware CPU optimization for Intel® Xeon® - production model serving, routing, and security built in.

  • Updated Aug 27, 2026
  • Shell

Production-ready platform for deploying OpenAI-compatible LLM inference on Kubernetes. A one-command installer for enterprise generative AI on Intel® Xeon®, with vLLM and other serving engines, secure Keycloak authentication, and built-in observability - from bare metal to a scalable AI endpoint.

  • Updated Aug 28, 2026
  • Shell

Enterprise-grade, high-performance LLM serving on Kubernetes. OpenAI-compatible inference powered by vLLM and OpenVINO™ Model Server behind a unified AI gateway, with NUMA-aware CPU optimization for Intel® Xeon® — production model serving, routing, and security built in.

  • Updated Aug 27, 2026
  • Shell

Production-ready platform for deploying OpenAI-compatible LLM inference on Kubernetes. A one-command installer for enterprise generative AI on Intel® Xeon®, with vLLM and other serving engines, secure Keycloak authentication, and built-in observability — from bare metal to a scalable AI endpoint.

  • Updated Aug 28, 2026
  • Shell

Improve this page

Add a description, image, and links to the model-serving topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the model-serving topic, visit your repo's landing page and select "manage topics."

Learn more