CrewFitnessSuite CRD
Overview
A CrewFitnessSuite runs a set of fitness scripts against a crew many times, collects the result of every iteration, and writes an XLSX report to the NATS Object Store bucket kubemoot_fitness_artifacts. The operator creates one CrewFitness per script and iteration, serially by default, and owns each one through an owner reference.
apiVersion: kubemoot.ai/v1alpha1
kind: CrewFitnessSuite
metadata:
name: homelab-baseline
namespace: crew-homelab-pilot
spec:
crewRef: homelab-pilot
description: "READ-query baseline across the homelab layers"
iterations: 15
concurrency: 1
scripts:
- testRef: gpu-utilization-live
configMapRef: homelab-pilot-fitness-tests
- testRef: node-count
configMapRef: homelab-pilot-fitness-tests
Spec
| Field | Type | Default | Description |
|---|---|---|---|
crewRef | string | required | Crew in the same namespace that the suite runs against. |
description | string | Purpose of the suite. Shown on the dashboard and in the XLSX Overview tab. | |
iterations | int | required, at least 1 | How many times each script runs. The suite creates len(scripts) x iterations runs. |
scripts | list | required, at least 1 | The scripts to run, in order. Each entry has a testRef and exactly one of testContent (inline script) or configMapRef (ConfigMap holding <testRef>.adl). |
concurrency | int | 1 | Maximum runs in flight at once. Keep it at 1 for baselines: parallel runs share the crew and its model providers, so they measure contention. |
perIterationTimeout | duration | 10m | Wall-clock cap for one iteration. |
artifactRetention | duration | 168h | Object Store TTL for the XLSX. |
purgeMemory | bool | true | Clear the crew’s working memory once before the first iteration. |
suspend | bool | false | Pause between iterations. See Pause, resume and stop. |
cancel | bool | false | Stop the suite. See Pause, resume and stop. |
rejudge | object | Set at creation, immutable. Judge an earlier run’s saved answers against this suite’s scripts instead of running iterations. suite names the source CrewFitnessSuite in the same namespace and runId its run. See Re-judge an earlier run. |
Status
| Field | Description |
|---|---|
phase | Pending, Running, Paused, Completed, Cancelled, Failed, or Error. |
runId | Identifies this execution. Used in iteration names and in the artifact key. |
startedAt | When scheduling began. |
completedAt | When the last iteration finished, or when the suite was cancelled. |
iterationsTotal | len(scripts) x iterations. It keeps the planned total after a cancel, so iterationsCompleted < iterationsTotal marks a partial run. |
iterationsCompleted | Iterations that reached a terminal phase, whatever the outcome. |
passed, failed, errored | iterationsCompleted broken down by outcome. |
artifactRef | Bucket, object key (<namespace>/<suite>/<runId>.xlsx), and size of the XLSX, set once it is written. |
error | Why the suite itself failed to execute (phase Error). |
rejudge | On a re-judge suite: the source run, how many transcripts were copied, and notInSource, the scenarios of this suite the source run has no answer for. |
scenarios | Per-scenario rollup of the finished iterations, in spec order. See Results in status. |
judge | Progress and scores of the deferred judge pass. See Results in status. |
conditions | On a re-judge suite, RejudgeSourceReady reports whether the source could be used. See Re-judge an earlier run. |
Phases
| Phase | Terminal | Meaning |
|---|---|---|
Pending | no | Created, not yet started. |
Running | no | Iterations are being scheduled or are in flight. |
Paused | no | spec.suspend is true and no iteration is in flight. Nothing runs against the crew. |
Completed | yes | Every iteration reached a terminal phase. Failed or errored iterations are outcomes in the counts, not a suite failure. |
Cancelled | yes | spec.cancel stopped the suite. Completed iterations are kept and the XLSX is marked partial. |
Failed, Error | yes | The suite itself could not execute, such as an invalid spec. |
Results in status
The operator writes a run’s results to the suite’s status, so any client with read access to crewfitnesssuites gets them through the Kubernetes API: kubectl, kmctl, CrewForge, or a script. The status outlives the per-iteration CrewFitness objects, which the operator removes after the run, and the object store retention of the transcripts and judge checkpoint.
status:
phase: Completed
runId: 7f3c2a1b
iterationsTotal: 30
iterationsCompleted: 30
passed: 26
failed: 3
errored: 1
scenarios:
- name: gpu-utilization-live
iterations: 15
passed: 14
failed: 1
meanDurationMs: 84210
assertionsPassed: 59
assertionsTotal: 60
- name: node-count
iterations: 15
passed: 12
failed: 2
errored: 1
meanDurationMs: 40112
assertionsPassed: 52
assertionsTotal: 60
judge:
phase: Complete
judged: 2
total: 2
mean: 44
zeros: 1
completedAt: "2026-10-02T14:31:07Z"
scores:
- scenario: gpu-utilization-live
score: 0
reason: all iterations gated (consensus failed or empty synthesis)
- scenario: node-count
score: 87
reason: Correct node count and roles; omits the control-plane taint the reference names.
artifactRef:
bucket: kubemoot_fitness_artifacts
objectKey: crew-homelab-pilot/homelab-baseline/7f3c2a1b.xlsx
sizeBytes: 48211
status.scenarios
One row per distinct testRef, in the order of spec.scripts. A script with no finished iteration yet has a row with only its name. The rows are refreshed on every reconcile of a running suite and are final once the suite is terminal.
| Field | Description |
|---|---|
name | The script’s testRef. |
iterations | Finished iterations of this scenario. |
passed, failed, errored | iterations broken down by outcome. |
meanDurationMs | Mean wall-clock duration of the finished iterations. |
assertionsPassed, assertionsTotal | Assertion results summed over the finished iterations. |
On a re-judge suite the rows come from the copied transcripts, with the durations of the source run.
status.judge
| Field | Description |
|---|---|
phase | Pending: the run has not finished, and the judge waits for it. Judging: the run finished and the judge is scoring it. Complete: every judgeable scenario has a score. Skipped: the suite was cancelled, so the judge does not run. |
judged | Scenarios scored so far. |
total | Scenarios the judge has to score: those with a DEFER assertion and at least one transcript. It is set when the pass starts. |
mean | Mean scenario score over the judged scenarios, 0 to 100, rounded. |
zeros | Judged scenarios that scored 0. A scenario whose every iteration failed its consensus floor scores 0 without a judge call. |
completedAt | When the operator recorded the pass as complete. |
scores | Each judged scenario in spec order: scenario, score (0 to 100, rounded), and reason, the judge’s rationale on one line, cut to 200 characters. |
The judge writes status.judge when its pass starts, after each scenario it scores, and when it finishes, so judged out of total shows progress while the pass runs. Once phase is Complete or Skipped the field is final: it is not rewritten, even after the judge checkpoint passes its object store retention. While judging, judged never goes down: if the checkpoint expires before the pass completes, status keeps the scores it recorded. status.judge is absent when the operator has no NATS Object Store configured, since there is no judge.
kubectl get crewfitnesssuites shows the judge phase and the mean score in the Judge and Quality columns.
What stays in the object store
The status carries summaries. The per-iteration transcripts (question, answer, events, assertion messages), the full judge reasons, the judge checkpoint, and the XLSX report stay in the NATS Object Store bucket kubemoot_fitness_artifacts, under <namespace>/<suite>/<runId>/, and status.artifactRef names the XLSX.
Size bound
Status is stored in etcd with the rest of the object, so both lists are capped at 300 entries (the first 300 distinct testRef values in spec order), scenario names at 253 bytes, and reasons at 200 characters. The API server rejects a longer list. judged, total, mean, and zeros always cover every scenario, so judged greater than the length of scores means the list was cut. A testRef longer than 253 bytes is cut in status, so keep names shorter than that to keep them distinct. For 300 scenarios with 30-character names and full reasons, the two lists take about 125 KiB; at the field maximums (253-byte names, reasons of 200 four-byte characters) they take under 450 KiB, well below the 1 MiB an object should stay under. A suite with inline testContent also stores its scripts in the spec, so keep very large script sets in a ConfigMap.
Pause, resume and stop
Two spec fields control a running suite. Edit them with kubectl patch, or use the Pause, Resume, and Stop buttons on the dashboard Fitness page.
Pause sets spec.suspend: true. The iteration in flight finishes and its result is kept; no new iteration starts. The phase stays Running until that iteration finishes, then becomes Paused. A Paused suite puts no load on the crew, which makes it safe to change the crew or the cluster between iterations.
Resume sets spec.suspend: false. The phase returns to Running and the suite continues at the next iteration it has not yet run. Completed iterations are not repeated.
Stop sets spec.cancel: true. The operator deletes any iteration in flight (its Job and pod follow through owner references), starts no new iteration, sets completedAt, and moves the suite to Cancelled. Stop works from Pending, Running, or Paused, takes precedence over suspend, and cannot be undone: clearing cancel afterwards does not restart the suite. Both fields are ignored once the suite is terminal.
# Pause
kubectl patch crewfitnesssuite homelab-baseline -n crew-homelab-pilot \
--type merge -p '{"spec":{"suspend":true}}'
# Resume
kubectl patch crewfitnesssuite homelab-baseline -n crew-homelab-pilot \
--type merge -p '{"spec":{"suspend":false}}'
# Stop
kubectl patch crewfitnesssuite homelab-baseline -n crew-homelab-pilot \
--type merge -p '{"spec":{"cancel":true}}'
The report of a cancelled suite
A cancelled suite still writes its XLSX. It holds only the iterations that finished before the stop; the iteration that was in flight is dropped. The Overview tab shows Status: Cancelled and a Partial row with how many iterations completed out of the planned total.
Deferred judging
Scores from DEFER assertions are produced by a judge pass that runs only after the suite reaches a terminal phase. A Running or Paused suite is never judged, so pausing cannot start judging early.
A Cancelled suite is not judged at all. Stop means no further work on the GPUs, and a judge pass over a partial run would score scenarios from fewer iterations than planned. The quality columns of a cancelled suite’s XLSX stay unjudged. To score a run, let it reach Completed.
Re-judge an earlier run
A judge score is only comparable to another judge score taken against the same references. When the references in the scenarios change (a corrected ground truth, a reworded reference), every earlier run was scored against text that no longer exists. spec.rejudge scores an earlier run’s saved answers against the current scenarios without asking the crew anything.
A re-judge suite is an ordinary CrewFitnessSuite whose scripts are the current scenarios, plus a rejudge block naming the source run:
apiVersion: kubemoot.ai/v1alpha1
kind: CrewFitnessSuite
metadata:
name: baseline-rejudged
namespace: crew-homelab-pilot
spec:
crewRef: homelab-pilot
description: "baseline answers judged against the current references"
iterations: 1
rejudge:
suite: baseline-n1 # the source CrewFitnessSuite, same namespace
runId: 001cd44d # the source's status.runId
scripts:
- testRef: gpu-utilization-all-rigs
testContent: |
...the current scenario text...
The simplest way to get the current scripts is to build the suite from the crew’s scenario directory, the same way a baseline is built, and add the rejudge block.
What happens
- Validate. The source suite must exist, its
status.runIdmust equalrejudge.runId, and its phase must beCompleted. Every script must resolve (inlinetestContent, or<testRef>.adlin itsconfigMapRef). - Copy. The operator reads the source run’s transcripts from the Object Store and matches them to this suite’s scripts by
testRef. Each matched transcript is copied into this suite’s own run, keeping the answer, the events, and the timing, and adding arejudgedFromfield with the source suite, run, and object key. - Re-evaluate the assertions. The copied transcript carries this suite’s assertions, not the source’s (see below).
- Complete. The suite moves from
Pendingstraight toCompleted.iterationsTotalis the number of transcripts copied, andpassedandfailedcome from the re-evaluated assertions. - Judge. The deferred judge pass then scores the copies exactly as it scores a normal run, against this suite’s
DEFERreferences, one scenario at a time with the same resumable checkpoint.status.judgereportsjudgedout oftotaluntil it finishes, and the XLSX is written once judging is complete.
The crew named in crewRef receives no question. iterations, concurrency, perIterationTimeout, purgeMemory, and suspend do not apply. The source run is not modified, and its own judge scores are not carried over.
Scenarios that do not match
- A scenario in this suite with no transcript in the source run (added since, or errored in the source) is listed in
status.rejudge.notInSourceand is not scored. - A scenario in the source run that this suite does not have is ignored.
Deterministic assertions
The fitness runner and the operator share one assertion engine, so the operator re-evaluates the deterministic ASSERT lines itself:
- An assertion whose text is identical to one the source run evaluated keeps the source result exactly as recorded.
- Any other assertion (new, or reworded) is evaluated against the stored answer and events. Its message ends with
(re-evaluated from the source transcript). A transcript is written only for an answered run, so the run state the runner had (POST returned 200, adoneevent arrived, no timeout) is recovered from the transcript itself. DEFERassertions carry this suite’s reference, which is what the judge reads.
When the source cannot be used
The suite goes to Error with status.error set and the RejudgeSourceReady condition False:
| Reason | Meaning |
|---|---|
InvalidRejudge | suite or runId is empty, suite names this suite, there are no scripts, or a testRef appears twice. |
SourceNotFound | No CrewFitnessSuite with that name in this namespace. Deleting a suite purges its transcripts, so a deleted run cannot be re-judged. |
RunIDMismatch | The source suite’s status.runId is not rejudge.runId. |
SourceNotCompleted | The source run is not Completed (still running, paused, cancelled, or failed). |
ScriptUnavailable | A script of this suite has no content: its ConfigMap or key is missing, or it sets both or neither source. A failure to reach the API server is retried, not reported. |
NoTranscripts | No scenario of this suite has a readable transcript in the source run (for example, the source transcripts passed their Object Store retention). |
NoArtifactStore | The operator has no NATS Object Store configured. |
On success the condition is True with reason SourceReady and a message naming the copy count and any scenarios not in the source. The rejudge block is set when the suite is created and cannot be added, changed, or removed afterwards; the API server rejects the edit. To re-judge again, create a new suite.
Scenarios are matched by testRef, so each testRef may appear only once in a re-judge suite. The source run’s transcripts are mapped to scenarios through the source suite’s spec.scripts; if those were edited after the run, the mapping follows the edited list.
With kmctl
kmctl has no re-judge flag or command; the rejudge block in the manifest is the whole interface.
kmctl fitness run -f baseline-rejudged.yaml -n crew-homelab-pilot # apply, wait for Completed
kmctl fitness get baseline-rejudged -n crew-homelab-pilot
kmctl fitness download baseline-rejudged -n crew-homelab-pilot # once judging has finished
Any suite, re-judge or not, reaches Completed before its judge pass finishes, so kmctl fitness run returns before the scores exist. A re-judge suite reaches Completed as soon as its transcripts are copied. Until judging ends, the XLSX served is provisional with quality 0; the operator rewrites it with the final scores. kmctl fitness get shows the judge progress and scores from status.judge. Do not pass --scenario for a re-judge suite: it creates a live CrewFitness that asks the crew the question.
Deleting a suite
Deleting the CrewFitnessSuite removes everything it owns. Its CrewFitness runs, their Jobs, and their pods are garbage-collected through owner references, and the operator’s finalizer (kubemoot.ai/fitness-artifacts) purges the run’s XLSX, transcripts, and judge scores from the Object Store before the resource goes. Delete a suite when you want the run and its results gone; stop it when you want to keep what completed.
Stopping a single CrewFitness
A standalone CrewFitness (one scenario, one run) has no stop field. Deleting it is the stop: its Job and pod are garbage-collected through owner references, and it has no partial result worth keeping.
kubectl delete crewfitness node-count-check -n crew-my-crew