Skip to content

Repository files navigation

self_hosted_inference_core logo

Hex package HexDocs MIT License

SelfHostedInferenceCore

self_hosted_inference_core is the service-runtime kernel for local and self-hosted inference backends.

It owns the runtime concerns that sit between raw process placement and backend-specific boot logic:

  • backend registration
  • runtime instance registration
  • startup-kind handling
  • readiness orchestration
  • health monitoring
  • lease and reuse semantics
  • endpoint publication
  • backend-to-consumer fit evaluation

It does not own transport mechanics or client protocol execution. execution_plane owns process placement, IO lifecycle, and raw process facts. req_llm remains the data-plane client after an endpoint has been resolved.

Runtime Stack

execution_plane
  -> self_hosted_inference_core
  -> concrete backend package or attach adapter
  -> req_llm consumers through EndpointDescriptor

That split keeps service lifecycle in the runtime stack and keeps request execution in the client layer.

Backends

Two backend shapes are now proved:

  • built-in attach adapter: SelfHostedInferenceCore.Ollama
  • concrete spawned backend package: llama_cpp_sdk
  • built-in service-mode simulation backend: SelfHostedInferenceCore.Simulation
  • supervised Crucible runtime workers: SelfHostedInferenceCore.CrucibleRuntime

SelfHostedInferenceCore.Ollama proves the first truthful management_mode: :externally_managed path. It attaches to an already running Ollama daemon, owns readiness and health interpretation above the transport seam, and publishes the same northbound endpoint contract used by the spawned path.

llama_cpp_sdk plugs into the kernel by implementing SelfHostedInferenceCore.Backend and owns:

  • llama-server boot-spec normalization
  • readiness and health probes
  • stop semantics for a spawned service
  • backend manifest publication
  • endpoint descriptor production

That keeps the kernel generic while proving both ownership shapes on real backends.

SelfHostedInferenceCore.Simulation is the built-in PRELIM service-mode simulation backend. It is selected by application configuration and a backend manifest, uses Execution Plane's :lower_simulation process transport, and publishes normal endpoint, readiness, health, and lease contracts without launching a real backend process. It is not selected through a public simulation: request option.

Startup Kinds

self_hosted_inference_core treats startup topology as an explicit part of the contract:

  • :spawned
    • BEAM-managed service lifecycle
    • maps to management_mode: :jido_managed
  • :attach_existing_service
    • externally managed daemon lifecycle
    • maps to management_mode: :externally_managed

Both paths use the same northbound endpoint and lease contracts. The kernel validates that backends keep startup kind, management mode, and substrate ownership truthful. It also rejects execution surfaces that are not declared in the backend manifest.

Governed Authority

Standalone self-hosted inference keeps the existing local ergonomics: Ollama.AttachSpec accepts direct endpoint config and endpoint auth. When no root URL is supplied, the attach default comes from config :self_hosted_inference_core, :ollama_root_url, ..., falling back to http://127.0.0.1:11434.

Governed self-hosted inference is selected with governed_authority: %SelfHostedInferenceCore.GovernedAuthority{} or equivalent attrs. In that mode the kernel and built-in Ollama attach adapter materialize endpoint identity, service identity, target posture, attach grant, credential refs, endpoint auth, backend options, execution surface, and redaction refs only from the authority packet.

When governed_authority is present, direct target-preference fields such as backend, startup_kind, execution_surface, backend_options, boot_spec, attach_spec, metadata, root_url, api_key, and headers are rejected. Direct Ollama attach fields such as root_url, base_url, model_identity, api_key, headers, ollama_http, execution_surface, and metadata are also rejected. Authority refs are copied into metadata, but materialized endpoint tokens are not copied into metadata, lease metadata, or instance ids.

Installation

Add the package to your dependency list:

def deps do
  [
    {:self_hosted_inference_core, "~> 0.2.0"}
  ]
end

Concrete backends register themselves against the kernel by implementing SelfHostedInferenceCore.Backend.

self_hosted_inference_core stays the service-runtime family kit. It consumes Execution Plane process facts for spawned services but keeps readiness, health, lease reuse, and endpoint publication semantics local to this repo.

See guides/backend_packages.md for how the kernel expects concrete backend packages to attach. See guides/ollama_attach.md for the built-in attached-local backend. See guides/crucible_runtime.md for supervised Crucible forward/generation workers.

Quick Start

Define a backend or attach adapter, register it, and ensure a northbound endpoint for a request:

alias SelfHostedInferenceCore.ConsumerManifest

:ok = SelfHostedInferenceCore.register_backend(MyBackend)

consumer =
  ConsumerManifest.new!(
    consumer: :jido_integration_req_llm,
    accepted_runtime_kinds: [:service],
    accepted_management_modes: [:jido_managed, :externally_managed],
    accepted_protocols: [:openai_chat_completions],
    required_capabilities: %{streaming?: true},
    optional_capabilities: %{},
    constraints: %{},
    metadata: %{adapter: :req_llm}
  )

request = %{
  request_id: "req-123",
  target_preference: %{
    target_class: "self_hosted_endpoint",
    backend: "my_backend",
    backend_options: %{model_identity: "demo-model"}
  }
}

context = %{
  run_id: "run-123",
  attempt_id: "run-123:1",
  boundary_ref: "boundary-123",
  observability: %{trace_id: "trace-123"}
}

{:ok, endpoint, fit_result} =
  SelfHostedInferenceCore.ensure_endpoint(
    request,
    consumer,
    context,
    owner_ref: "run-123",
    ttl_ms: 30_000
  )

endpoint.base_url
endpoint.lease_ref
fit_result.reason

See examples/README.md for runnable demos covering both :spawned and :attach_existing_service.

HexDocs

HexDocs includes:

  • architecture and stack-boundary guidance
  • built-in ollama attach guidance
  • concrete backend package guidance
  • the northbound endpoint contract used by jido_integration
  • runtime registry, lease reuse, and endpoint publication semantics
  • startup-kind guidance for spawned and attached services
  • runnable examples

License

Released under the MIT License. See LICENSE.

Crucible Runtime Status

Status: provider-boundary-in-progress.

SelfHostedInferenceCore.CrucibleRuntime is provider-neutral. The kernel owns supervision, readiness, leases, health snapshots, and request delegation. Concrete provider packages own model loading, preflight, signal extraction, trace writing, and generation.

The provider-bound live gate shape is:

CRUCIBLE_LIVE_MODEL=true mix run examples/crucible_runtime_live.exs -- \
  --provider-module SelfHostedInferenceBumblebee.CrucibleProvider \
  --model hf-internal-testing/tiny-random-gpt2 \
  --backend binary

Live model proof artifacts are provider-owned. The core package only requires that providers return canonical %Crucible.ForwardTrace{} values and explicit generation results or structured errors.

About

Core Elixir primitives for building reliable self-hosted inference clients, provider adapters, transport boundaries, and operational controls for private AI runtimes across local, edge, and dedicated infrastructure.

Topics

Resources

Stars

Watchers

Forks

Releases

Packages

Contributors

Languages