Models in Kubemoot

This document covers the model layer end to end: the CRDs (ModelProvider, Model, EmbeddingModel), how Agents reach a (model, provider) binding via the scheduler, the load-on-demand principle, the model catalog, and GPU compatibility for a reference deployment.

For the framework behind these choices - why tool-calling fidelity outranks coding benchmarks, how to budget VRAM, and how to match model size to discussion role - see Choosing a Model.

CRD Hierarchy

ModelProvider (where - a model server endpoint)
    │
    ├── Model (what - labeled inference model on a provider)
    │
    └── EmbeddingModel (what - embedding model for RAGSource indexing)

Agents declare capabilities; the scheduler matches them to a Model whose labels satisfy a CrewSchedulingPolicy’s require / prefer selectors. See Scheduler for the full algorithm.

ModelProvider

Declares an inference endpoint. The operator discovers GPU capacity (VRAM, loaded models, agent count) via DCGM metrics and Ollama APIs.

Spec

FieldTypeRequiredDescription
typestringYesOnly ollama is supported today. Any other value, including openai and anthropic, is rejected at admission: those types are not supported yet.
endpointstringFor ollamaAPI endpoint URL
secretRefstringNoReserved for future provider types; not used by ollama

Status

FieldSourceDescription
readyOperator probeEndpoint reachable + API responding
capacity.vramTotalMiBDCGM PrometheusTotal GPU VRAM
capacity.vramUsedMiBOllama /api/psVRAM consumed by loaded models
capacity.gpuModelDCGM metricse.g., “NVIDIA GeForce RTX 5090”
capacity.maxParallelPod env OLLAMA_NUM_PARALLELConcurrent request slots
capacity.contextLengthPod env OLLAMA_CONTEXT_LENGTHPer-request context (tokens) each slot gives a model; unknown when the engine chooses its own default
capacity.loadedModels[].contextLengthOllama /api/psContext the loaded model actually runs with; takes precedence over capacity.contextLength
capacity.agentCountOperatorAgents currently bound to this provider
capacity.loadedModelsOllama /api/psModels currently resident
capacity.availableModelsOllama /api/tagsModels downloaded on this provider
capacity.nodeNameKubernetesNode hosting the Ollama pod
capacity.lastProbedOperatorWhen capacity was last discovered

A ModelProvider stored with any type other than ollama (one created before the API server rejected them) reports ready: false with reason Unsupported and the same message the API server gives; Models and EmbeddingModels on it report the same message.

The scheduler refuses to place a call whose prompt exceeds a provider’s context for the model, because the engine would silently drop the oldest messages instead of rejecting it (see Scheduler). On Ollama the total KV cache scales with OLLAMA_NUM_PARALLEL times OLLAMA_CONTEXT_LENGTH, so a GPU trades parallel slots against per-request context.

Discovery needs the scheduler enabled and, for VRAM, DCGM metrics in Prometheus. A provider the operator cannot measure (a CPU Ollama, or a GPU host without DCGM) declares its budget instead:

spec:
  scheduling:
    memoryMiB: 6144   # VRAM on a GPU host; host RAM the server may use on a CPU host

A discovered VRAM total always wins over the declaration. With neither, agents refuse cold loads on that provider because the fit gate has no budget to check against.

Example

apiVersion: kubemoot.ai/v1alpha1
kind: ModelProvider
metadata:
  name: ollama-a
spec:
  type: ollama
  endpoint: http://ollama.ollama-a:11434
---
apiVersion: kubemoot.ai/v1alpha1
kind: ModelProvider
metadata:
  name: ollama-b
spec:
  type: ollama
  endpoint: http://ollama.ollama-b:11434

kubectl get mdlp (short name: mdlp).

Model

Declares a specific inference model on a provider, labeled so the scheduler can select it. The operator pulls the model if needed and tracks its lifecycle.

Spec

FieldTypeRequiredDescription
providerRefstringYesName of the ModelProvider CR
modelstringYesModel identifier on the provider (e.g., qwen3:32b, gpt-4o)
vramMibint32For GPU providersCapacity request - used by the scheduler’s filter step
contextLengthintNoOverride maximum context window
quantizationstringNoQuantization method (e.g., q4_K_M, q8_0)

Labels (open string set, selected by CrewSchedulingPolicy)

LabelPurposeExamples
familyModel familyqwen3, mistral, glm, deepseek-r1, phi3
paramsParameter count"3B", "8B", "14B", "24B", "32B"
capability/<name>What the model can docapability/tool-calling: "true", capability/reasoning: "true"
contextWindowMax context tokens"32768", "131072", "200000"
latencyClassLatency profilelow, medium, high

Status

FieldDescription
statePending, Pulling, Available, Loaded, Error
readyModel ready for inference
endpointInference endpoint (from provider)
modelInfoSize, parameters, family, quantization, context length, format, digest

Example

apiVersion: kubemoot.ai/v1alpha1
kind: Model
metadata:
  name: qwen3-32b
  labels:
    family: qwen3
    params: "32B"
    capability/tool-calling: "true"
    capability/reasoning: "true"
    contextWindow: "32768"
    latencyClass: high
spec:
  model: qwen3:32b
  providerRef: ollama-a
  vramMib: 20480
  quantization: q4_K_M
---
apiVersion: kubemoot.ai/v1alpha1
kind: Model
metadata:
  name: qwen3-8b
  labels:
    family: qwen3
    params: "8B"
    capability/tool-calling: "true"
    contextWindow: "131072"
    latencyClass: low
spec:
  model: qwen3:8b
  providerRef: ollama-b
  vramMib: 5120
  quantization: q4_K_M

kubectl get mdl (short name: mdl).

EmbeddingModel

Declares an embedding model for RAGSource indexing. Not consumed by chat agents; RAGSources reference EmbeddingModels via spec.embeddingModelRef. It is a separate resource from Model because its lifecycle and consumers differ (see the RAGSource design notes).

Spec

FieldTypeDefaultDescription
providerRefstring(required)Name of the ModelProvider CR
modelstring(required)Model name (e.g., nomic-embed-text, text-embedding-3-small)
dimensionsint32768Vector dimension size
batchSizeint3232Max texts per embedding request

Status

FieldDescription
statePending, Pulling, Available, Error
readyModel ready for embedding
endpointEmbedding API endpoint (from provider)
modelInfoDimensions, max input tokens, family

Example

apiVersion: kubemoot.ai/v1alpha1
kind: EmbeddingModel
metadata:
  name: nomic-embed
spec:
  model: nomic-embed-text
  providerRef: ollama-a

kubectl get emb (short name: emb).

How Agents Reach a Model

The Agent CR does not name a model or a provider. It declares capabilities:

apiVersion: kubemoot.ai/v1alpha1
kind: Agent
metadata:
  name: kubectl-agent
spec:
  capabilities: [tool-calling, kubernetes]
  discussRole: tooler

The crew’s CrewSchedulingPolicy declares which models satisfy each discussion phase via Kubernetes-style label selectors:

apiVersion: kubemoot.ai/v1alpha1
kind: CrewSchedulingPolicy
metadata:
  name: homelab-pilot-default
  namespace: crew-homelab-pilot
spec:
  crewRef: homelab-pilot
  rules:
    - phase: mulling
      require:
        matchLabels: { capability/tool-calling: "true" }
        matchExpressions:
          - { key: params, operator: In, values: ["14B", "32B"] }
      prefer:
        - weight: 100
          selector:
            matchLabels: { family: qwen3, params: "32B" }
    - phase: triage
      require:
        matchLabels: { capability/tool-calling: "true" }
      prefer:
        - weight: 100
          selector:
            matchLabels: { latencyClass: low }

The scheduler computes a (model, provider, endpoint) per agent via filter → score → bind and templates the result into the agent’s Deployment as KUBEMOOT_MODEL_MODEL, KUBEMOOT_MODEL_ENDPOINT, KUBEMOOT_TRIAGE_MODEL_MODEL_ID, KUBEMOOT_TRIAGE_MODEL_ENDPOINT. See Scheduler for the algorithm.

application.properties wiring

The agent-runtime is Quarkus + LangChain4j. The quarkus.langchain4j.ollama.chat-model.model-id property is BUILD_AND_RUN_TIME_FIXED and cannot be overridden via env vars at runtime. The operator writes resolved values to application.properties in the {agent}-policy ConfigMap, mounted at /app/config/application.properties (SmallRye ordinal 260, overrides in-jar ordinal 250 at startup).

quarkus.langchain4j.ollama.chat-model.model-id=qwen3:32b
quarkus.langchain4j.ollama.base-url=http://ollama.ollama-a:11434
quarkus.langchain4j.ollama.chat-model.temperature=0.3
quarkus.langchain4j.ollama.chat-model.num-predict=2048

temperature and num-predict come from Agent.spec.temperature and Agent.spec.maxTokens.

Load on Demand - We Do Not Preload

Kubemoot loads a model into provider VRAM when an agent asks for it, not before. There is no preload field, no warm-up CRD, no eager-loading hook. This is intentional design, not a missing feature.

The principle

Load a model when an agent needs it. Do not tie up a GPU with a model it’s not needed.

GPU VRAM is shared capacity. Pinning model X in advance for a guess at what crew A will use starves model Y for whichever agent actually fires next. Any system that hand-declares preload intent is making predictions about future agent demand that the scheduler can - and should - make at request time.

What handles the cold-load case

When the first inference of the session hits a cold provider, Ollama loads the model from disk to VRAM. This takes 30-60s for a 30B-class model on an RTX 5090 (and longer on the 4090). After that, the model is resident.

Two mechanisms make this cost acceptable:

  1. OLLAMA_KEEP_ALIVE=1h - once loaded, the model stays in VRAM for an hour of inactivity. Every subsequent query within the session is warm.
  2. CrewSchedulingPolicy prefer weights - direct the scheduler toward models the provider is likely to already have resident (via image-locality scoring) and toward models that fit the phase’s latency profile.

The cold-load tax is paid once per (model, provider) per session. If a user complains that “the first response was slow,” the answer is not to add preload. The answer is to verify that subsequent queries are fast (they will be), and to consider session-shaping if the cold cost is unacceptable for the workload.

When to push back on preload

Future requests in the form “can we just preload model X on GPU Y” should be re-framed:

  1. What problem is the preload solving? (Usually: “first query is slow.”)
  2. Is the slowness on the first query only, or on every query? (If only first: KEEP_ALIVE handles the rest - accept the cold cost.)
  3. If genuinely needed, address at the infrastructure layer: increase KEEP_ALIVE, tune the scheduler, or admit that the workload pattern is wrong for shared GPU capacity.
  4. Pin VRAM only as a last resort, and never via a chart - manual curl /api/generate keep_alive=-1 is fine for one-off debugging.

The architectural commitment: the orchestrator decides what to load. The chart declares what’s possible (Model CRs with labels), not what’s resident.

Model Catalog

Sizes below are approximate and move with each quantization release. Treat them as planning figures and validate the real footprint against ModelProvider.status.capacity after a cold load.

Tier 1: Frontier Models (Cloud-Routed, Not Locally Hostable)

These mixture-of-experts families are one to two orders of magnitude beyond a single consumer GPU. On local runtimes they generally exist only as cloud-routed tags, which resolve to the vendor’s hosted infrastructure rather than your hardware. Reaching them is a deliberate decision to move inference off the premises.

ModelOrgTotal ParamsActiveTool CallingOllama TagSelf-Host Footprint
Kimi K2 family (K2.6, K2.7-Code, K3)Moonshot AI1T MoE~32BNativekimi-k2.6:cloud, kimi-k3, kimi-k2.7-code~630 GB full weights; ~240 GB at 1.8-bit dynamic quant; roughly 4x H200 to self-host
GLM-5.2Z.aiMoE-Nativeglm-5.2:cloud~223 GB at 1-bit; 256 GB+ at 2-bit
DeepSeek V4 ProDeepSeekMoE-Nativedeepseek-v4-pro:cloudData-centre class
MiniMax M3MiniMaxMoE-Yesminimax-m3Data-centre class

For scale: four H200 class accelerators run roughly $120,000 to $160,000 to buy, or about $4 to $18 per hour to rent. That is the price of admission to this tier, which is why the practical homelab question is which 20B to 35B class model calls tools best.

These tiers are not reachable from Kubemoot today: hosted providers are on the Roadmap.

Tier 2: Locally Hostable Models (20B-35B Class)

The working range for a single GPU with 24 GB to 48 GB of memory, which is roughly a $5,000 card. Comfortable on a 32 GB card at Q4; workable on a 24 GB card at Q4 with less KV cache headroom.

MoE caveat: mixture-of-experts models load all expert weights into VRAM, so size follows total params. The active-param count buys speed, not VRAM savings. Validate the real cold-load footprint against provider capacity, not the active-param number.

ModelOrgParamsTool CallingOllama TagApprox Q4Notes
Qwen3.6 35B-A3BAlibaba35B MoE (3B active)Yesqwen3.6:35b~20 GBStrong reported agentic tool use for its class. MoE speed at 35B footprint.
Qwen3.6 27BAlibaba27B denseYesqwen3.6:27b~17 GBLong context. Fits a 24 GB card with headroom.
Laguna XS 2.1-33B MoE (3B active)Yeslaguna-xs-2.1~20 GBAgentic-coding tuned.
Nemotron 3NVIDIA33BNative (agentic)nemotron3:33b~20 GBAgentic-tuned. NVIDIA-Open license.
Gemma 4 26B / 31BGoogle26B / 31BYesgemma4:26b, gemma4:31b~16 / ~19 GBLarge install base, permissive. Reported tool use is weaker than the Qwen3.6 class; prefer elsewhere for Toolers.
Granite 4.1 30BIBM30BYesgranite4.1:30b~18 GBBuilt around function calling. Permissive license.
LFM2 24BLiquid AI24BYeslfm2:24b~14 GBFits both reference GPUs at Q4.
MiniMax M2.7MiniMax-Yesminimax-m2.7-Validate footprint before committing a phase to it.
Qwen3 32BAlibaba32B denseYesqwen3:32b~20 GBLong-standing default in the reference crew. Superseded by the Qwen3.6 class on tool use.

Tier 3: Small Models (3B-14B Class)

Suitable for Tooler agents, low-priority discussion participants, or a 24 GB card with room to spare.

ModelOrgParamsTool CallingOllama TagApprox Q4Notes
Qwen3.5 9BAlibaba9BYesqwen3.5:9b~6 GBCurrent small Qwen line.
Granite 4.1 8BIBM8BYesgranite4.1:8b~5 GBFunction-calling focus at Tooler size. Permissive license.
Granite 4.1 3BIBM3BYesgranite4.1:3b~2 GBLarge co-schedule headroom.
LFM2.5 8BLiquid AI8BYeslfm2.5:8b~5 GBLow-latency Tooler candidate.
Gemma 4 12BGoogle12BYesgemma4:12b~8 GBPermissive, widely available.
Qwen3 14BAlibaba14BYesqwen3:14b~9 GBDual-mode (thinking / non-thinking).
Qwen3 8BAlibaba8BYesqwen3:8b~5 GBDual-mode. Long-standing Tooler default in the reference crew.

A model with a toggleable thinking mode can regress badly in a Tooler slot when reasoning is disabled, because it stops planning which tool to call. Confirm the mode the runtime actually sends rather than assuming the default.

GPU Compatibility

Measured on the reference homelab: an RTX 5090 node (ollama-a) and an RTX 4090 node (ollama-b). The quantization, VRAM, and headroom figures below are specific to that hardware; treat them as a worked example of the budgeting method in Choosing a Model, not as universal numbers for a different card.

RTX 5090 - 32 GB VRAM (ollama-a)

ModelQuantizationVRAMKV HeadroomFits?Recommended Labels
Qwen3.6:35b-a3bQ4_K_M~20 GB~12 GBYesfamily=qwen3.6, params=35B, capability/tool-calling=true, capability/reasoning=true, latencyClass=medium
Qwen3.6:27bQ4_K_M~17 GB~15 GBYesfamily=qwen3.6, params=27B, capability/tool-calling=true, latencyClass=medium
Qwen3:32bQ4_K_M~20 GB~12 GBYesfamily=qwen3, params=32B, capability/tool-calling=true, capability/reasoning=true, latencyClass=high
Laguna XS 2.1Q4_K_M~20 GB~12 GBYesfamily=laguna, params=33B, capability/tool-calling=true, latencyClass=medium
Nemotron3:33bQ4_K_M~20 GB~12 GBYesfamily=nemotron3, params=33B, capability/tool-calling=true, latencyClass=high
Granite4.1:30bQ4_K_M~18 GB~14 GBYesfamily=granite4.1, params=30B, capability/tool-calling=true, latencyClass=medium
Gemma4:31bQ4_K_M~19 GB~13 GBYesfamily=gemma4, params=31B, capability/tool-calling=true, latencyClass=medium
LFM2:24bQ4_K_M~14 GB~18 GBYesfamily=lfm2, params=24B, capability/tool-calling=true, latencyClass=medium
Qwen3:14bQ4_K_M~9 GB~23 GBYesfamily=qwen3, params=14B, capability/tool-calling=true, latencyClass=medium
Qwen3:8bQ4_K_M~5 GB~27 GBYesfamily=qwen3, params=8B, capability/tool-calling=true, latencyClass=low
Any 32B classQ8_0~35 GBn/aNoExceeds VRAM

RTX 4090 - 24 GB VRAM (ollama-b)

Constraint: num_parallel=1 required - num_parallel=2 causes page cache OOM preventing model reload.

ModelQuantizationVRAMKV HeadroomFits?Recommended Labels
Qwen3.6:27bQ4_K_M~17 GB~7 GBYesfamily=qwen3.6, params=27B, capability/tool-calling=true, latencyClass=medium
LFM2:24bQ4_K_M~14 GB~10 GBYesfamily=lfm2, params=24B, capability/tool-calling=true, latencyClass=medium
Gemma4:12bQ4_K_M~8 GB~16 GBYesfamily=gemma4, params=12B, capability/tool-calling=true, latencyClass=low
Qwen3:14bQ4_K_M~9 GB~15 GBYesfamily=qwen3, params=14B, capability/tool-calling=true, latencyClass=medium
Qwen3.5:9bQ4_K_M~6 GB~18 GBYesfamily=qwen3.5, params=9B, capability/tool-calling=true, latencyClass=low
Granite4.1:8bQ4_K_M~5 GB~19 GBYesfamily=granite4.1, params=8B, capability/tool-calling=true, latencyClass=low
LFM2.5:8bQ4_K_M~5 GB~19 GBYesfamily=lfm2.5, params=8B, capability/tool-calling=true, latencyClass=low
Qwen3:8bQ4_K_M~5 GB~19 GBYesfamily=qwen3, params=8B, capability/tool-calling=true, latencyClass=low
Granite4.1:3bQ4_K_M~2 GB~22 GBYesfamily=granite4.1, params=3B, capability/tool-calling=true, latencyClass=low
Any 32B classQ4_K_M~20 GB~4 GBTightLoads, but too little KV cache for long discussions
Any 24B classQ8_0~25 GBn/aNoExceeds VRAM

Choosing Among These

The selection framework lives in Choosing a Model: rank on tool-call fidelity rather than coding benchmarks, budget VRAM with KV cache headroom, size mixture-of-experts models by total parameters, match model size to discussion role, and decide by fitness measurement rather than by leaderboard.

Prefer permissive licenses (Apache-2.0, MIT, and comparable) and avoid non-commercial terms.

Troubleshooting

Model stuck in Pending

The controller checks if the model exists on the provider. For Ollama, it calls the /api/tags endpoint. If the model isn’t pulled yet, pull it manually:

kubectl exec -n ollama deploy/ollama -- ollama pull qwen3:32b

Agent not using the expected model

Check the agent’s scheduling status:

kubectl get agent <name> -o jsonpath='{.status.phase}{"\n"}{.status.message}{"\n"}'

If the phase is Unschedulable, the CrewSchedulingPolicy require selectors didn’t match any feasible Model. Check Model labels:

kubectl get models --show-labels

Either relax the require selectors or add a Model that matches them.

Temperature and other inference settings

temperature and maxTokens live on the Agent CR. They’re written to application.properties by the operator and cannot be overridden via environment variables (Quarkus BUILD_AND_RUN_TIME_FIXED constraint). Update the Agent CR and wait for the operator to regenerate the ConfigMap.