Quickstart

Try Kubemoot on a laptop, no GPU required, in one script.

This is a laptop trial: one script takes an empty Kubernetes cluster to a real crew answering a real question, on CPU, with no GPU required. It runs the same protocol as a production deployment - a coordinator convening a Tooler, signals exchanged over NATS, an answer synthesized - just smaller and slower, because it runs on CPU with a 1-3B model instead of a GPU with a reasoning-sized one. See Installation for the GPU-backed deployment this trial is a preview of, and the profile table below for exactly what’s smaller and slower.

Run it

git clone https://github.com/kubemoot/kubemoot.git && cd kubemoot
kind create cluster --name kubemoot     # or any cluster your kubectl points at
./quickstart/quickstart.sh

Needs kubectl, helm, and python3. The script needs a node with 2 CPUs and 8 GB free and takes several minutes on a laptop; the model pull and the first discussion dominate. Release images are amd64 only today (see the Roadmap), so on an arm64 machine such as Apple Silicon the images run under emulation or fail to start with an exec format error. Every wait polls a Kubernetes status field rather than sleeping a fixed time, so the script finishes as soon as the cluster is actually ready, not on a guessed clock.

What it installs

StepWhatWhy
1NATS with JetStreamEvery discussion is carried over NATS; the operator bootstraps its streams.
2Ollama on CPU, pulling qwen2.5:1.5bA small model that runs anywhere, with no GPU.
3The operator Helm chart, minimal profile, with the dashboard onThe dashboard has no login, so it gets no route; the script prints the kubectl port-forward command that opens it. Admission webhooks, Grafana dashboards, the internal MCP servers, the NATS MCP server, and the verify runner are off.
4A ModelProvider pointing at that OllamaThe endpoint the scheduler binds agents to at inference time.
5The hello crew: a coordinator plus one ToolerThe smallest crew that actually deliberates; its prompts are ADL rules you can kubectl apply and change.
6One question through the crew’s discussion gatewayThe answer comes back as the discussion’s synthesis, streamed over Server-Sent Events.

The CPU trial profile

This is what’s different about a CPU trial versus a GPU deployment, so you can judge Kubemoot’s speed and a crew’s size by the right yardstick:

CPU trial (this script)GPU deployment
Model size1-3B parameters (qwen2.5:1.5b here)8B-32B+ per role, sized to VRAM
Agents per crewTwo (a coordinator plus one Tooler)As many as the crew’s domain needs
Answer latencySlower, dominated by CPU inferenceFaster, bounded by GPU inference
Concurrent discussionsOne at a timeSeveral, bounded by GPU capacity

What just happened

You did three things, all declaratively: installed an operator, applied a crew, and asked it a question. Under the hood the coordinator selected the one Tooler that fits the question, the Tooler answered from its own knowledge, and the answer was settled by the crew rather than asserted by one model - see consensus signal for what that means. This is the same protocol a GPU deployment runs; only the model size, the crew size, and the clock are different.

Troubleshooting

  • The answer says no agent contributed. The 1.5B model in the trial sometimes stands aside. Rerun the question; the script is idempotent.
  • The NATS pod stays Pending. NATS keeps its stream on a PersistentVolumeClaim, so the cluster needs a default StorageClass. kind provides one.
  • Ollama restarts or is killed. The node needs 8 GB free; on a smaller node the kernel kills Ollama while it loads the model.
  • A pod shows ImagePullBackOff. Check that the node can reach ghcr.io, and that it runs on amd64.

Then

kubectl get agents,models,crew -n hello
kubectl edit promptmodule hello-facts-system -n hello   # change a rule, ask again

Ask another question by rerunning with QUESTION="..." EXPECT="" set, or with kmctl conversation ask hello "..." -n hello. The script is idempotent: rerunning it against the same cluster upgrades in place rather than failing.

CI runs this exact script on a fresh kind cluster against every release’s images, so if it breaks, the build is red before you see it. Full details, environment variables, and what the script does not cover (RAG sources, MCP tool servers, GPU scheduling across providers, a published dashboard) live in quickstart/README.md.

Next

  • Installation - install the operator on your own cluster with a GPU-backed model provider, for real crew work rather than a trial.
  • Crews & Agents - the model you just ran.
  • The starter crew - scaffold a read-only guide to a namespace with kmctl create, install it, ask it, and change one of its rules.
  • Build a Crew - author your own, from scratch or from the packaged Homelab Pilot chart.
  • CrewForge - develop the crew from VS Code: create, deploy, ask, and test it in the editor.
  • Define Fitness Functions - measure whether the crew actually does its job.