CrewFitnessSuite CRD

Overview

A CrewFitnessSuite runs a set of fitness scripts against a crew many times, collects the result of every iteration, and writes an XLSX report to the NATS Object Store bucket kubemoot_fitness_artifacts. The operator creates one CrewFitness per script and iteration, serially by default, and owns each one through an owner reference.

apiVersion: kubemoot.ai/v1alpha1
kind: CrewFitnessSuite
metadata:
  name: homelab-baseline
  namespace: crew-homelab-pilot
spec:
  crewRef: homelab-pilot
  description: "READ-query baseline across the homelab layers"
  iterations: 15
  concurrency: 1
  scripts:
    - testRef: gpu-utilization-live
      configMapRef: homelab-pilot-fitness-tests
    - testRef: node-count
      configMapRef: homelab-pilot-fitness-tests

Spec

FieldTypeDefaultDescription
crewRefstringrequiredCrew in the same namespace that the suite runs against.
descriptionstringPurpose of the suite. Shown on the dashboard and in the XLSX Overview tab.
iterationsintrequired, at least 1How many times each script runs. The suite creates len(scripts) x iterations runs.
scriptslistrequired, at least 1The scripts to run, in order. Each entry has a testRef and exactly one of testContent (inline script) or configMapRef (ConfigMap holding <testRef>.adl).
concurrencyint1Maximum runs in flight at once. Keep it at 1 for baselines: parallel runs share the crew and its model providers, so they measure contention.
perIterationTimeoutduration10mWall-clock cap for one iteration.
artifactRetentionduration168hObject Store TTL for the XLSX.
purgeMemorybooltrueClear the crew’s working memory once before the first iteration.
suspendboolfalsePause between iterations. See Pause, resume and stop.
cancelboolfalseStop the suite. See Pause, resume and stop.
rejudgeobjectSet at creation, immutable. Judge an earlier run’s saved answers against this suite’s scripts instead of running iterations. suite names the source CrewFitnessSuite in the same namespace and runId its run. See Re-judge an earlier run.

Status

FieldDescription
phasePending, Running, Paused, Completed, Cancelled, Failed, or Error.
runIdIdentifies this execution. Used in iteration names and in the artifact key.
startedAtWhen scheduling began.
completedAtWhen the last iteration finished, or when the suite was cancelled.
iterationsTotallen(scripts) x iterations. It keeps the planned total after a cancel, so iterationsCompleted < iterationsTotal marks a partial run.
iterationsCompletedIterations that reached a terminal phase, whatever the outcome.
passed, failed, errorediterationsCompleted broken down by outcome.
artifactRefBucket, object key (<namespace>/<suite>/<runId>.xlsx), and size of the XLSX, set once it is written.
errorWhy the suite itself failed to execute (phase Error).
rejudgeOn a re-judge suite: the source run, how many transcripts were copied, and notInSource, the scenarios of this suite the source run has no answer for.
scenariosPer-scenario rollup of the finished iterations, in spec order. See Results in status.
judgeProgress and scores of the deferred judge pass. See Results in status.
conditionsOn a re-judge suite, RejudgeSourceReady reports whether the source could be used. See Re-judge an earlier run.

Phases

PhaseTerminalMeaning
PendingnoCreated, not yet started.
RunningnoIterations are being scheduled or are in flight.
Pausednospec.suspend is true and no iteration is in flight. Nothing runs against the crew.
CompletedyesEvery iteration reached a terminal phase. Failed or errored iterations are outcomes in the counts, not a suite failure.
Cancelledyesspec.cancel stopped the suite. Completed iterations are kept and the XLSX is marked partial.
Failed, ErroryesThe suite itself could not execute, such as an invalid spec.

Results in status

The operator writes a run’s results to the suite’s status, so any client with read access to crewfitnesssuites gets them through the Kubernetes API: kubectl, kmctl, CrewForge, or a script. The status outlives the per-iteration CrewFitness objects, which the operator removes after the run, and the object store retention of the transcripts and judge checkpoint.

status:
  phase: Completed
  runId: 7f3c2a1b
  iterationsTotal: 30
  iterationsCompleted: 30
  passed: 26
  failed: 3
  errored: 1
  scenarios:
    - name: gpu-utilization-live
      iterations: 15
      passed: 14
      failed: 1
      meanDurationMs: 84210
      assertionsPassed: 59
      assertionsTotal: 60
    - name: node-count
      iterations: 15
      passed: 12
      failed: 2
      errored: 1
      meanDurationMs: 40112
      assertionsPassed: 52
      assertionsTotal: 60
  judge:
    phase: Complete
    judged: 2
    total: 2
    mean: 44
    zeros: 1
    completedAt: "2026-10-02T14:31:07Z"
    scores:
      - scenario: gpu-utilization-live
        score: 0
        reason: all iterations gated (consensus failed or empty synthesis)
      - scenario: node-count
        score: 87
        reason: Correct node count and roles; omits the control-plane taint the reference names.
  artifactRef:
    bucket: kubemoot_fitness_artifacts
    objectKey: crew-homelab-pilot/homelab-baseline/7f3c2a1b.xlsx
    sizeBytes: 48211

status.scenarios

One row per distinct testRef, in the order of spec.scripts. A script with no finished iteration yet has a row with only its name. The rows are refreshed on every reconcile of a running suite and are final once the suite is terminal.

FieldDescription
nameThe script’s testRef.
iterationsFinished iterations of this scenario.
passed, failed, errorediterations broken down by outcome.
meanDurationMsMean wall-clock duration of the finished iterations.
assertionsPassed, assertionsTotalAssertion results summed over the finished iterations.

On a re-judge suite the rows come from the copied transcripts, with the durations of the source run.

status.judge

FieldDescription
phasePending: the run has not finished, and the judge waits for it. Judging: the run finished and the judge is scoring it. Complete: every judgeable scenario has a score. Skipped: the suite was cancelled, so the judge does not run.
judgedScenarios scored so far.
totalScenarios the judge has to score: those with a DEFER assertion and at least one transcript. It is set when the pass starts.
meanMean scenario score over the judged scenarios, 0 to 100, rounded.
zerosJudged scenarios that scored 0. A scenario whose every iteration failed its consensus floor scores 0 without a judge call.
completedAtWhen the operator recorded the pass as complete.
scoresEach judged scenario in spec order: scenario, score (0 to 100, rounded), and reason, the judge’s rationale on one line, cut to 200 characters.

The judge writes status.judge when its pass starts, after each scenario it scores, and when it finishes, so judged out of total shows progress while the pass runs. Once phase is Complete or Skipped the field is final: it is not rewritten, even after the judge checkpoint passes its object store retention. While judging, judged never goes down: if the checkpoint expires before the pass completes, status keeps the scores it recorded. status.judge is absent when the operator has no NATS Object Store configured, since there is no judge.

kubectl get crewfitnesssuites shows the judge phase and the mean score in the Judge and Quality columns.

What stays in the object store

The status carries summaries. The per-iteration transcripts (question, answer, events, assertion messages), the full judge reasons, the judge checkpoint, and the XLSX report stay in the NATS Object Store bucket kubemoot_fitness_artifacts, under <namespace>/<suite>/<runId>/, and status.artifactRef names the XLSX.

Size bound

Status is stored in etcd with the rest of the object, so both lists are capped at 300 entries (the first 300 distinct testRef values in spec order), scenario names at 253 bytes, and reasons at 200 characters. The API server rejects a longer list. judged, total, mean, and zeros always cover every scenario, so judged greater than the length of scores means the list was cut. A testRef longer than 253 bytes is cut in status, so keep names shorter than that to keep them distinct. For 300 scenarios with 30-character names and full reasons, the two lists take about 125 KiB; at the field maximums (253-byte names, reasons of 200 four-byte characters) they take under 450 KiB, well below the 1 MiB an object should stay under. A suite with inline testContent also stores its scripts in the spec, so keep very large script sets in a ConfigMap.

Pause, resume and stop

Two spec fields control a running suite. Edit them with kubectl patch, or use the Pause, Resume, and Stop buttons on the dashboard Fitness page.

Pause sets spec.suspend: true. The iteration in flight finishes and its result is kept; no new iteration starts. The phase stays Running until that iteration finishes, then becomes Paused. A Paused suite puts no load on the crew, which makes it safe to change the crew or the cluster between iterations.

Resume sets spec.suspend: false. The phase returns to Running and the suite continues at the next iteration it has not yet run. Completed iterations are not repeated.

Stop sets spec.cancel: true. The operator deletes any iteration in flight (its Job and pod follow through owner references), starts no new iteration, sets completedAt, and moves the suite to Cancelled. Stop works from Pending, Running, or Paused, takes precedence over suspend, and cannot be undone: clearing cancel afterwards does not restart the suite. Both fields are ignored once the suite is terminal.

# Pause
kubectl patch crewfitnesssuite homelab-baseline -n crew-homelab-pilot \
  --type merge -p '{"spec":{"suspend":true}}'

# Resume
kubectl patch crewfitnesssuite homelab-baseline -n crew-homelab-pilot \
  --type merge -p '{"spec":{"suspend":false}}'

# Stop
kubectl patch crewfitnesssuite homelab-baseline -n crew-homelab-pilot \
  --type merge -p '{"spec":{"cancel":true}}'

The report of a cancelled suite

A cancelled suite still writes its XLSX. It holds only the iterations that finished before the stop; the iteration that was in flight is dropped. The Overview tab shows Status: Cancelled and a Partial row with how many iterations completed out of the planned total.

Deferred judging

Scores from DEFER assertions are produced by a judge pass that runs only after the suite reaches a terminal phase. A Running or Paused suite is never judged, so pausing cannot start judging early.

A Cancelled suite is not judged at all. Stop means no further work on the GPUs, and a judge pass over a partial run would score scenarios from fewer iterations than planned. The quality columns of a cancelled suite’s XLSX stay unjudged. To score a run, let it reach Completed.

Re-judge an earlier run

A judge score is only comparable to another judge score taken against the same references. When the references in the scenarios change (a corrected ground truth, a reworded reference), every earlier run was scored against text that no longer exists. spec.rejudge scores an earlier run’s saved answers against the current scenarios without asking the crew anything.

A re-judge suite is an ordinary CrewFitnessSuite whose scripts are the current scenarios, plus a rejudge block naming the source run:

apiVersion: kubemoot.ai/v1alpha1
kind: CrewFitnessSuite
metadata:
  name: baseline-rejudged
  namespace: crew-homelab-pilot
spec:
  crewRef: homelab-pilot
  description: "baseline answers judged against the current references"
  iterations: 1
  rejudge:
    suite: baseline-n1        # the source CrewFitnessSuite, same namespace
    runId: 001cd44d           # the source's status.runId
  scripts:
    - testRef: gpu-utilization-all-rigs
      testContent: |
        ...the current scenario text...

The simplest way to get the current scripts is to build the suite from the crew’s scenario directory, the same way a baseline is built, and add the rejudge block.

What happens

  1. Validate. The source suite must exist, its status.runId must equal rejudge.runId, and its phase must be Completed. Every script must resolve (inline testContent, or <testRef>.adl in its configMapRef).
  2. Copy. The operator reads the source run’s transcripts from the Object Store and matches them to this suite’s scripts by testRef. Each matched transcript is copied into this suite’s own run, keeping the answer, the events, and the timing, and adding a rejudgedFrom field with the source suite, run, and object key.
  3. Re-evaluate the assertions. The copied transcript carries this suite’s assertions, not the source’s (see below).
  4. Complete. The suite moves from Pending straight to Completed. iterationsTotal is the number of transcripts copied, and passed and failed come from the re-evaluated assertions.
  5. Judge. The deferred judge pass then scores the copies exactly as it scores a normal run, against this suite’s DEFER references, one scenario at a time with the same resumable checkpoint. status.judge reports judged out of total until it finishes, and the XLSX is written once judging is complete.

The crew named in crewRef receives no question. iterations, concurrency, perIterationTimeout, purgeMemory, and suspend do not apply. The source run is not modified, and its own judge scores are not carried over.

Scenarios that do not match

  • A scenario in this suite with no transcript in the source run (added since, or errored in the source) is listed in status.rejudge.notInSource and is not scored.
  • A scenario in the source run that this suite does not have is ignored.

Deterministic assertions

The fitness runner and the operator share one assertion engine, so the operator re-evaluates the deterministic ASSERT lines itself:

  • An assertion whose text is identical to one the source run evaluated keeps the source result exactly as recorded.
  • Any other assertion (new, or reworded) is evaluated against the stored answer and events. Its message ends with (re-evaluated from the source transcript). A transcript is written only for an answered run, so the run state the runner had (POST returned 200, a done event arrived, no timeout) is recovered from the transcript itself.
  • DEFER assertions carry this suite’s reference, which is what the judge reads.

When the source cannot be used

The suite goes to Error with status.error set and the RejudgeSourceReady condition False:

ReasonMeaning
InvalidRejudgesuite or runId is empty, suite names this suite, there are no scripts, or a testRef appears twice.
SourceNotFoundNo CrewFitnessSuite with that name in this namespace. Deleting a suite purges its transcripts, so a deleted run cannot be re-judged.
RunIDMismatchThe source suite’s status.runId is not rejudge.runId.
SourceNotCompletedThe source run is not Completed (still running, paused, cancelled, or failed).
ScriptUnavailableA script of this suite has no content: its ConfigMap or key is missing, or it sets both or neither source. A failure to reach the API server is retried, not reported.
NoTranscriptsNo scenario of this suite has a readable transcript in the source run (for example, the source transcripts passed their Object Store retention).
NoArtifactStoreThe operator has no NATS Object Store configured.

On success the condition is True with reason SourceReady and a message naming the copy count and any scenarios not in the source. The rejudge block is set when the suite is created and cannot be added, changed, or removed afterwards; the API server rejects the edit. To re-judge again, create a new suite.

Scenarios are matched by testRef, so each testRef may appear only once in a re-judge suite. The source run’s transcripts are mapped to scenarios through the source suite’s spec.scripts; if those were edited after the run, the mapping follows the edited list.

With kmctl

kmctl has no re-judge flag or command; the rejudge block in the manifest is the whole interface.

kmctl fitness run -f baseline-rejudged.yaml -n crew-homelab-pilot   # apply, wait for Completed
kmctl fitness get baseline-rejudged -n crew-homelab-pilot
kmctl fitness download baseline-rejudged -n crew-homelab-pilot      # once judging has finished

Any suite, re-judge or not, reaches Completed before its judge pass finishes, so kmctl fitness run returns before the scores exist. A re-judge suite reaches Completed as soon as its transcripts are copied. Until judging ends, the XLSX served is provisional with quality 0; the operator rewrites it with the final scores. kmctl fitness get shows the judge progress and scores from status.judge. Do not pass --scenario for a re-judge suite: it creates a live CrewFitness that asks the crew the question.

Deleting a suite

Deleting the CrewFitnessSuite removes everything it owns. Its CrewFitness runs, their Jobs, and their pods are garbage-collected through owner references, and the operator’s finalizer (kubemoot.ai/fitness-artifacts) purges the run’s XLSX, transcripts, and judge scores from the Object Store before the resource goes. Delete a suite when you want the run and its results gone; stop it when you want to keep what completed.

Stopping a single CrewFitness

A standalone CrewFitness (one scenario, one run) has no stop field. Deleting it is the stop: its Job and pod are garbage-collected through owner references, and it has no partial result worth keeping.

kubectl delete crewfitness node-count-check -n crew-my-crew