Define Fitness Functions
A crew’s behavior is only as trustworthy as your ability to check it. Kubemoot implements architectural fitness functions as Kubernetes resources: you write a short scenario that asks the crew a real question and states what a good answer must satisfy, apply it, and the operator runs the crew live and scores each assertion pass or fail.
The shape of a scenario
A fitness scenario is ADL pseudo-code - a question plus assertions:
DESCRIPTION Verify the crew answers a basic infrastructure question.
DEFINE CONST QUESTION AS "How many nodes are in the cluster?"
DEFINE CONST MAX_DURATION AS 120 seconds
ASSERT(discussion completes within MAX_DURATION)
ASSERT(at least 1 specialist agrees)
ASSERT(coordinator produces synthesis)
ASSERT(synthesis CONTAINS reference to a node count)
Common assertion patterns include: the discussion endpoint returns 200; a named SSE event is emitted (optionally within a deadline); the discussion completes within a time budget; at least N specialists (Toolers) agree (or, for an out-of-scope smoke test, zero agree); the coordinator produces a synthesis; and the synthesis contains or does not contain specific terms. In an assertion, write “specialist” for a Tooler: the runner’s agree-floor assertion matches that word, even though the docs call them Toolers.
Run one scenario
Store scenarios as a ConfigMap alongside the crew and declare a CrewFitness resource
that references one:
apiVersion: kubemoot.ai/v1alpha1
kind: CrewFitness
metadata:
name: node-count-check
namespace: crew-my-crew
spec:
crewRef: my-crew
testRef: node-count # key in the ConfigMap
configMapRef: my-crew-fitness-tests
ttl: 1h # auto-delete after completion
The operator validates the crew has a discussion endpoint, runs a job that asks the
question and streams the discussion, evaluates every ASSERT against the collected
signals, and writes per-assertion pass/fail into the resource’s status. With ttl set,
the CrewFitness cleans itself up after it finishes.
kubectl get crewfitness -n crew-my-crew -w
kubectl describe crewfitness node-count-check -n crew-my-crew # per-assertion results
Run many - suites and baselines
To measure stability rather than a single pass, a CrewFitnessSuite runs a scenario
(or several) many times and writes a spreadsheet of results - pass rates and p50/p90
timings per scenario - so you can see a crew hold steady or drift across releases.
A long suite can be paused, resumed, or stopped while it runs, from the dashboard Fitness
page or with kubectl patch:
- Pause (
spec.suspend: true): the iteration in flight finishes and is kept, no new one starts, and the suite readsPausedonce nothing is running. - Resume (
spec.suspend: false): the suite continues at the next iteration it has not run. - Stop (
spec.cancel: true): the iteration in flight is ended, the suite readsCancelled, and the spreadsheet is written for the iterations that completed, marked partial. A stopped suite is not quality-judged.
kubectl patch crewfitnesssuite my-baseline -n crew-my-crew --type merge -p '{"spec":{"suspend":true}}'
When you change a scenario’s reference, earlier scores were taken against the old text.
A suite with spec.rejudge scores an earlier run’s saved answers against the current
scenarios without asking the crew again, so old and new runs compare on the same footing.
To stop a single CrewFitness, delete it. See the
CrewFitnessSuite reference for every field and phase.
Holistic by design
These are holistic fitness functions: each run exercises the whole crew at once - coordinator, Toolers, Analysts, tool gateway, message bus, and live model inference - against real, measurable behavior, not a mocked unit. That is what makes “does the crew still do its job?” an objective, repeatable check instead of a transcript you eyeball.
For the deeper treatment, see Kubemoot Crew Fitness Functions.