Install the Operator
This page installs the Kubemoot operator: what a cluster needs before it runs, and how to install it on your own cluster - the GPU-backed deployment Kubemoot is built for. If you want to try the protocol first without a GPU, the Quickstart runs a smaller CPU-only trial with its own script and its own limits, described below.
The operator is one part of the ecosystem. The others install separately, each from its own page:
- kmctl, the command-line tool: kmctl User Guide.
- Crews, packaged as Helm charts, including the reference crew: Crews and Pilot.
- CrewForge, the VS Code extension: CrewForge.
- Integrations that reach a crew from elsewhere, such as Claude Code: Integrations.
Prerequisites
- A Kubernetes cluster (v1.30+) and
kubectlconfigured against it. Cluster-admin is required for the initial install, since Kubemoot installs CRDs and cluster-scoped RBAC. - Helm (v3.8+, for OCI charts). The operator ships only as a Helm chart.
- A model provider - at least one GPU-backed inference endpoint that serves the models your crew will use. Kubemoot is built for Ollama today; one or more GPUs are the intended target. A crew composes capability from several small models, so a single modest GPU is enough to start. The GPU must be visible to the worker node, on bare metal or passed through to a VM; running a model server on it is covered in Put a model server on your GPU nodes below. Ollama can be installed in-cluster with the community Ollama Helm chart.
- NATS JetStream - the message bus crews deliberate over. It can run in-cluster; the operator publishes discussion and event traffic to it. Install it with the official NATS Helm chart with JetStream enabled.
- A vector store (pgvector) - only if a crew uses RAG knowledge sources. Not required for a discussion-only crew. See pgvector for PostgreSQL with vector search.
CPU trial vs. GPU deployment
Kubemoot does not require a GPU to run. The Quickstart runs on a node with 2 CPUs and 8 GB free and no accelerator, and its CPU trial profile table lists what differs from a GPU deployment: smaller models, a smaller crew, and slower answers. The rest of this page installs the GPU-backed profile.
Release images are amd64 only today (see the Roadmap). On an arm64 cluster or laptop they run under emulation or fail to start with an exec format error.
Install the operator
The operator ships as the kubemoot-operator Helm chart. The chart installs the CRDs,
the controller deployment, and its RBAC.
helm upgrade --install kubemoot-operator \
<chart-reference> \
--namespace kubemoot --create-namespace
Replace <chart-reference> with the chart source. From a checkout of the repository
you can install the bundled chart directly:
helm upgrade --install kubemoot-operator \
operator/chart/kubemoot-operator \
--namespace kubemoot --create-namespace
The published chart is oci://ghcr.io/kubemoot/charts/kubemoot-operator; pass it as
the chart reference with --version for a specific release.
Optional: the dashboard
The operator chart carries the dashboard as an optional subchart. It is off by default because the dashboard has no login and can purge NATS streams and delete models. To install it with the operator:
helm upgrade --install kubemoot-operator \
oci://ghcr.io/kubemoot/charts/kubemoot-operator \
--namespace kubemoot --create-namespace \
--set dashboard.enabled=true
It has no route; reach it with kubectl port-forward, as the
dashboard page shows. The dashboard chart is also published
by itself if you prefer a separate release.
Optional: GitOps with Flux
If you already manage your cluster with GitOps, you can skip the helm command above.
In a GitOps setup the operator is reconciled from the chart by a Flux HelmRelease
pointed at an OCIRepository, rather than installed imperatively. Set
install.crds: Create and upgrade.crds: CreateReplace so CRDs are managed with the
release. This is how the reference deployment runs Kubemoot.
Verify
kubectl get pods -n kubemoot
kubectl get crds | grep kubemoot.ai
You should see the operator pod Running and the Kubemoot CRDs registered
(crews, agents, models, modelproviders, mcpservers, mcpgateways,
ragsources, promptmodules, and the cluster-scoped kubemootconfig).
Put a model server on your GPU nodes
Kubemoot does not install Ollama, and it does not pick GPU nodes for you. The seam is deliberate: placement is yours, capacity is discovered.
- You run a model server per GPU worker node. A common topology is a worker node
per GPU, and a cluster scales by adding such nodes. For every GPU node, deploy its own
Ollama (or other provider) as an ordinary workload pinned to that node by node
affinity, requesting
nvidia.com/gpu, with the runtime class your cluster uses for NVIDIA. Each server becomes one provider with its GPU’s VRAM, and the scheduler places models across the providers, loading and evicting them on demand. The CPU trial runs a single Deployment with no GPU at all. - You declare a
ModelProviderper model server. One manifest for each: the type, that server’s Service URL, a weight, and, when the host cannot be measured, a memory budget. Two GPU nodes, two model servers, two ModelProviders. This is the only static declaration the operator needs. - The operator discovers the rest. It probes the endpoint for available and loaded
models, follows the Service to its backing pod, records that pod’s node, reads
OLLAMA_NUM_PARALLEL, and, when a Prometheus endpoint is configured, queries the DCGM metrics for that node to learn the GPU model and its total and used VRAM. Without DCGM,spec.scheduling.memoryMiBis the budget. - Agents choose a provider per call. At each inference call the runtime fits the model it needs against every ready provider’s discovered headroom and picks one. So “sized to VRAM” in the table above means sized to the VRAM discovered on the provider that serves the call, not to the largest GPU in the cluster: a 32B model needs a provider whose GPU holds it, and smaller roles land on the smaller GPU.
apiVersion: kubemoot.ai/v1alpha1
kind: ModelProvider
metadata:
name: ollama-gpu-a # the Ollama pinned to GPU worker node A
namespace: kubemoot
spec:
type: ollama
endpoint: http://ollama.ollama-a:11434
weight: 100
---
apiVersion: kubemoot.ai/v1alpha1
kind: ModelProvider
metadata:
name: ollama-gpu-b # the Ollama pinned to GPU worker node B
namespace: kubemoot
spec:
type: ollama
endpoint: http://ollama.ollama-b:11434
weight: 100
The full field reference, discovery sources, and the budget rules are in Models.
What you’ll see before a model provider is ready
The operator, a crew, and its agents can all exist as Kubernetes resources before a
ModelProvider is reachable. Nothing crashes; nothing answers either. Knowing what
that state looks like saves you from wondering whether the install is broken:
ModelProviderreportsstatus.phase: Failedandstatus.ready: falsewith a message such asFailed to connect to Ollama: ...until the operator can reach the endpoint and parse a version response. Once it can,status.phasebecomesReady.Agentreportsstatus.phase: Unschedulableandstatus.ready: falsewith messageno feasible Model for mulling phase(or a more specific scheduling reason) when noModelon aReadyModelProvidersatisfies its declared capabilities. An unschedulable agent gets no Deployment at all - there is no crash-looping pod to find, because none was created. The operator keeps retrying on a backoff until a provider appears.Crewfollows its agents. It reportsstatus.phase: Pendinguntil its coordinatorAgentisRunning, andDegradedwithstatus.ready: falsewhen the coordinator isUnschedulableor when every specialist is, with a message naming the agents concerned (coordinator coordinator is unschedulable: no feasible Model for mulling phase, orno specialist can be scheduled: k8s-advisor, ...). A crew with a running coordinator and at least one schedulable specialist isReady, and its message lists any specialists still unschedulable.kubectl get crews -Atherefore tells you whether a crew can answer before you ask it.
Once a ModelProvider becomes Ready and the operator reconciles, agents move to
Running and pick up their Deployments without you having to do anything.
Next
- Quickstart - deploy a crew and ask it a question.
- Build a Crew - author your own crew.