Modelplane Modelplane docs

ModelService Custom Resource

This document is for an unreleased version of Modelplane.

This document applies to the Modelplane main branch and not to the latest release v0.4.

A ModelService is one model as a caller sees it: a stable name that resolves to whichever ModelEndpoint should serve the next request. The endpoints behind it can be replicas Modelplane runs, models bought from a provider, or both, in more than one region. A caller reaches it by naming it as the model in an ordinary OpenAI or Anthropic request to any InferenceGateway that serves it.

Concept guide: Expose a Model →

#Metadata

API version
modelplane.ai/v1alpha1
Kind
ModelService
Scope
Namespaced
Short names
ms

#Example

Manifest
apiVersion: modelplane.ai/v1alpha1
kind: ModelService
metadata:
  name: qwen-72b
  namespace: ml-team
  labels:
    # Matched by an InferenceGateway's serviceSelector. Your label, under your
    # own prefix.
    example.org/region: eu
spec:
  endpoints:
    # Entries at the same priority share traffic by weight, so this pair is a
    # 90/10 canary across two deployments. name is a stable handle, unique
    # within the service.
    - name: stable
      priority: 0
      weight: 90
      selector:
        matchLabels:
          modelplane.ai/deployment: qwen-72b
    - name: canary
      priority: 0
      weight: 10
      selector:
        matchLabels:
          modelplane.ai/deployment: qwen-72b-next
    # Priority 1 takes traffic as priority 0 loses healthy endpoints, which
    # makes this provider a failover for the deployments above.
    - name: together
      priority: 1
      selector:
        matchLabels:
          modelplane.ai/endpoint: together-qwen-72b

#Spec

# endpoints required object[] 1–32 items
# name required string 1–63 chars

A stable name for this entry, unique within the service. For example stable, canary, or a provider’s name.

pattern: ^[a-z0-9]([-a-z0-9]*[a-z0-9])?$

# priority optional integer ≤ 63

Lower is preferred. Entries at the same priority share traffic by weight. A higher-numbered entry takes a growing share of traffic as lower-numbered ones lose healthy endpoints, and takes over entirely once they have none. A failed request is retried, on another endpoint at the same priority if there is one and then at the next, and each attempt gets that endpoint’s own model name, credential and path. Retrying is only possible until the first byte reaches the caller, because after that the tokens are already sent, so a backend that dies mid-stream truncates the response instead.

format: int32

# selector required object

Selects ModelEndpoints in this ModelService’s namespace. Scope a service to a region by selecting only endpoints in it; Modelplane stamps an InferenceCluster’s labels onto every endpoint composed there, so the region is declared once on the cluster.

# matchLabels required map[string]string
# weight optional integer 1–1000000 default: 1

Share of traffic for this entry relative to the other entries at the same priority, spread as evenly as possible across the endpoints it matches. A pair of entries weighted 90 and 10 is a canary. At least 1. A weight of 0 doesn’t deprioritise a backend, it drops it from the gateway’s load assignment entirely, which is indistinguishable from removing the entry and easy to mistake for parking it. Remove the entry instead.

format: int32

# timeouts optional object

How long a gateway waits on this service’s endpoints. The right values depend on the model and on whether callers stream, so set them from its observed response times.

# idle optional string default: 60s

How long an endpoint may send nothing. Before its first byte, the gateway abandons it, counts it as failed, and retries the request, on another endpoint if there is one. After that the stream is cut short. A streamed response sends its first byte after prefill, so for streaming callers this bounds time to first token and every gap between chunks. A non-streamed response sends nothing until it’s complete, so if any caller doesn’t stream, set this at least as long as request, or to 0s to disable it.

pattern: ^([0-9]{1,5}(h|m|s|ms)){1,4}$

# request optional string default: 300s

How long a request may take end to end, including retries. Set it above the longest response you expect: prefill plus the maximum output tokens at the model’s decode rate.

pattern: ^([0-9]{1,5}(h|m|s|ms)){1,4}$

#Status

# conditions optional object[]
# model optional string

The name a caller passes as the request’s model. Namespaced, so two services can’t collide.

# routes optional object

Counts of the ModelRoutes this service composes, one per gateway that serves it. ready is how many are carrying traffic; total is how many gateways serve the service. Per-gateway detail, including each gateway’s address, is on the ModelRoutes themselves: kubectl get modelroutes -l modelplane.ai/service=.

# ready optional integer

format: int32

# total optional integer

format: int32