<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Concepts on Kubemoot</title><link>https://kubemoot.org/docs/concepts/</link><description>Recent content in Concepts on Kubemoot</description><generator>Hugo</generator><language>en</language><atom:link href="https://kubemoot.org/docs/concepts/index.xml" rel="self" type="application/rss+xml"/><item><title>The Moot - Consensus Model</title><link>https://kubemoot.org/docs/concepts/consensus-model/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://kubemoot.org/docs/concepts/consensus-model/</guid><description>&lt;p&gt;Kubemoot&amp;rsquo;s defining idea is that an answer is reached by &lt;strong&gt;deliberation among
agents&lt;/strong&gt;, not asserted by a single model. This page explains the goal, the
signal vocabulary, the discussion phases, and, importantly, that &lt;em&gt;how&lt;/em&gt; a crew
organizes its deliberation is configurable. For the units of interaction the
deliberation sits inside - conversations, turns, and the thread each discussion
runs on - see &lt;a href="../conversations-turns-and-threads/"&gt;Conversations, Turns &amp;amp; Threads&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id="the-goal"&gt;The goal&lt;/h2&gt;
&lt;p&gt;Settle a question through structured deliberation rather than by one authority. A
crew of agents each contributes what it knows; the result is composed from the
crew, so a single model being confidently wrong does not, by itself, become the
answer.&lt;/p&gt;</description></item><item><title>Conversations, Turns &amp; Threads</title><link>https://kubemoot.org/docs/concepts/conversations-turns-and-threads/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://kubemoot.org/docs/concepts/conversations-turns-and-threads/</guid><description>&lt;p&gt;When you ask a crew a question, four words describe what happens, and they nest.
The docs and the dashboard use them with specific meanings, so this page defines
each one and how it relates to the others.&lt;/p&gt;
&lt;h2 id="the-layers"&gt;The layers&lt;/h2&gt;
&lt;pre tabindex="0"&gt;&lt;code&gt; Conversation durable session with a crew (conversationId)
 |
 +-- Turn one question and the answer settled on it
 | |
 | +-- Thread where that turn&amp;#39;s discussion runs (threadId)
 | |
 | +-- Discussion the deliberation on the thread (phases + signals)
 |
 +-- Turn ... later questions, each on its own thread
&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;A conversation holds many turns; each turn runs on its own thread; the discussion
is what happens on that thread. The two innermost words name the same layer from
different angles, which is explained below.&lt;/p&gt;</description></item><item><title>Crews &amp; Agents</title><link>https://kubemoot.org/docs/concepts/crews-and-agents/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://kubemoot.org/docs/concepts/crews-and-agents/</guid><description>&lt;p&gt;A &lt;strong&gt;crew&lt;/strong&gt; is a team of agents that deliberate together; an &lt;strong&gt;agent&lt;/strong&gt; is one
participant in that team. Both are Kubernetes resources you declare and version like
any other workload.&lt;/p&gt;
&lt;h2 id="crew"&gt;Crew&lt;/h2&gt;
&lt;p&gt;A &lt;code&gt;Crew&lt;/code&gt; groups a set of &lt;code&gt;Agent&lt;/code&gt;s - joined by the &lt;code&gt;kubemoot.ai/crew&lt;/code&gt; label - into one
collaborating team, and configures crew-wide concerns: the discussion gateway (the
HTTP entry point external clients talk to) and the crew&amp;rsquo;s &lt;strong&gt;working memory&lt;/strong&gt;. Agents
are the workers; the crew is the team-level context they share.&lt;/p&gt;</description></item><item><title>The Table and the Harnesses</title><link>https://kubemoot.org/docs/concepts/the-table-and-the-harnesses/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://kubemoot.org/docs/concepts/the-table-and-the-harnesses/</guid><description>&lt;p&gt;A harness is everything around a model: the guides that steer an agent before it acts
and the sensors that let it correct itself afterwards, in the sense of Martin Fowler&amp;rsquo;s
&lt;a href="https://martinfowler.com/articles/harness-engineering.html"&gt;Harness Engineering&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;A crew is a harness made of harnesses. Each agent has its own loop: its prompt, its
tools, its retries, its memory. The moot is the harness one level up, and its guides
and sensors act between agents:&lt;/p&gt;</description></item><item><title>ADL, the Architecture Definition Language, for agents</title><link>https://kubemoot.org/docs/concepts/agent-definition-language/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://kubemoot.org/docs/concepts/agent-definition-language/</guid><description>&lt;p&gt;Kubemoot defines what an agent does in &lt;strong&gt;ADL&lt;/strong&gt;, a small, structured language of
behavioral rules. Every agent prompt is ADL stored in a &lt;code&gt;PromptModule&lt;/code&gt; resource, never
free prose in an Agent spec, and the same language expresses a crew&amp;rsquo;s executable fitness
checks. ADL is the surface where a crew&amp;rsquo;s behavior is authored, reviewed, and versioned.&lt;/p&gt;
&lt;h2 id="what-adl-is-and-where-it-comes-from"&gt;What ADL is, and where it comes from&lt;/h2&gt;
&lt;p&gt;ADL is the &lt;strong&gt;Architecture Definition Language&lt;/strong&gt;, a pseudo-code introduced by Mark
Richards for describing and governing the structure of a software system, and paired
with the fitness-function idea from Richards and Neal Ford&amp;rsquo;s evolutionary-architecture
work (see &lt;a href="https://learning.oreilly.com/library/view/architecture-as-code/9798341640368/"&gt;&lt;em&gt;Architecture as Code&lt;/em&gt;&lt;/a&gt; by Mark Richards, Neal Ford, and
Jonathan Johnson, and Richards&amp;rsquo; ADL reference at developertoarchitect.com). In its original form ADL defines a system&amp;rsquo;s parts (&lt;code&gt;DEFINE SYSTEM&lt;/code&gt;, &lt;code&gt;DEFINE DOMAIN&lt;/code&gt;, &lt;code&gt;DEFINE COMPONENT&lt;/code&gt;) and asserts the rules between them
(&lt;code&gt;ASSERT&lt;/code&gt;, &lt;code&gt;FOREACH ... CONTAINED WITHIN&lt;/code&gt;), making architecture a machine-readable
artifact and enforcing it with executable checks that fail when an implementation
drifts from its intended design. Kubemoot applies that same language, unchanged in
form, to a different subject: instead of governing how a system&amp;rsquo;s components may
depend on each other, its rules govern how an agent behaves. This is a partnership with
ADL&amp;rsquo;s authors and their notation, not a fork or a renaming.&lt;/p&gt;</description></item><item><title>Signals &amp; Protocol</title><link>https://kubemoot.org/docs/concepts/signals-and-protocol/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://kubemoot.org/docs/concepts/signals-and-protocol/</guid><description>&lt;p&gt;Agents in a crew never call each other directly. They publish to a message bus
(NATS JetStream), and every contribution carries a &lt;strong&gt;signal&lt;/strong&gt; that says how it relates
to the emerging answer. The signals plus a phase model are the whole protocol - there
is no central controller issuing commands.&lt;/p&gt;
&lt;h2 id="consensus-signals"&gt;Consensus signals&lt;/h2&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;Signal&lt;/th&gt;
 &lt;th&gt;Meaning&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;&lt;code&gt;agree&lt;/code&gt;&lt;/td&gt;
 &lt;td&gt;This contribution supports the emerging answer.&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;&lt;code&gt;concern&lt;/code&gt;&lt;/td&gt;
 &lt;td&gt;A reservation that should be weighed before the crew settles.&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;&lt;code&gt;stand_aside&lt;/code&gt;&lt;/td&gt;
 &lt;td&gt;No relevant contribution; abstain without blocking. With &lt;code&gt;metadata.reason&lt;/code&gt; set to &lt;code&gt;gpu-busy&lt;/code&gt;, &lt;code&gt;model-too-large&lt;/code&gt;, or &lt;code&gt;prompt-too-large&lt;/code&gt;, the agent was willing but could not get a GPU (see below).&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;&lt;code&gt;block&lt;/code&gt;&lt;/td&gt;
 &lt;td&gt;A strong objection that should stop the answer as it stands.&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;&lt;code&gt;failure&lt;/code&gt;&lt;/td&gt;
 &lt;td&gt;The agent tried and could not complete - surfaced as a first-class signal, not hidden behind silence.&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Treating &lt;strong&gt;failure as signal&lt;/strong&gt; is deliberate. An agent that fails a tool call or
cannot reach a source says so, so the coordinator can route around it, rather than
standing aside silently and letting the crew mistake &amp;ldquo;no answer&amp;rdquo; for &amp;ldquo;no objection.&amp;rdquo;&lt;/p&gt;</description></item><item><title>Discussion Artifact Store</title><link>https://kubemoot.org/docs/concepts/discussion-artifact-store/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://kubemoot.org/docs/concepts/discussion-artifact-store/</guid><description>&lt;p&gt;When a tooler fetches data from a live system, the result can be large: a full
namespace inventory, a Prometheus time-series dump, a JSON response carrying
hundreds of rows. The NATS discussion bus and the LLM context window were not
designed to carry datasets. Dumping bulk data onto the bus truncates it, floods
downstream agents&amp;rsquo; context, and degrades the quality of reasoning that follows.&lt;/p&gt;
&lt;p&gt;Kubemoot passes large outputs &lt;strong&gt;by reference&lt;/strong&gt;. The data lands in an object store;
the discussion message carries only a small reference and a preview. Agents that
need the data fetch a targeted slice; most can answer from the preview alone.&lt;/p&gt;</description></item><item><title>Models &amp; Scheduling</title><link>https://kubemoot.org/docs/concepts/models-and-scheduling/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://kubemoot.org/docs/concepts/models-and-scheduling/</guid><description>&lt;p&gt;Kubemoot composes capability from &lt;strong&gt;many small models&lt;/strong&gt; rather than one large one,
and it binds an agent to a model &lt;strong&gt;as late as possible&lt;/strong&gt;. Two ideas make that work:
loose model coupling and just-in-time scheduling.&lt;/p&gt;
&lt;h2 id="loose-model-coupling"&gt;Loose model coupling&lt;/h2&gt;
&lt;p&gt;An &lt;code&gt;Agent&lt;/code&gt; spec names no model. It declares abstract &lt;strong&gt;capabilities&lt;/strong&gt; -
&lt;code&gt;tool-calling&lt;/code&gt;, &lt;code&gt;reasoning&lt;/code&gt;, &lt;code&gt;kubernetes&lt;/code&gt;, and so on. Concrete models are separate
resources: a &lt;code&gt;Model&lt;/code&gt; (an LLM, by capability labels, not a hardcoded name) served by a
&lt;code&gt;ModelProvider&lt;/code&gt; (a GPU-backed inference endpoint such as Ollama on a particular GPU).&lt;/p&gt;</description></item><item><title>Choosing a Model</title><link>https://kubemoot.org/docs/concepts/choosing-a-model/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://kubemoot.org/docs/concepts/choosing-a-model/</guid><description>&lt;p&gt;Kubemoot runs on the GPUs you have. A crew is typically twenty or more agents sharing a
small number of cards, so choosing a model is a capacity decision as much as a quality
one. This page is the decision framework. For the CRD fields, the current model catalog,
and the per-GPU fit matrix, see &lt;a href="../../reference/models/"&gt;Models in Kubemoot&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Because an &lt;code&gt;Agent&lt;/code&gt; names no model, changing your mind is cheap. Adopting a different
model is a &lt;code&gt;Model&lt;/code&gt; CR and a label, not an edit to every agent in the crew. That makes
model choice a thing you &lt;strong&gt;measure and revise&lt;/strong&gt;, not a thing you get right up front.&lt;/p&gt;</description></item><item><title>MCP Tools</title><link>https://kubemoot.org/docs/concepts/mcp-tools/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://kubemoot.org/docs/concepts/mcp-tools/</guid><description>&lt;p&gt;An agent&amp;rsquo;s knowledge comes from its model and its RAG sources; its &lt;em&gt;abilities&lt;/em&gt; - to
read a cluster, query a database, search the web - come from &lt;strong&gt;tools&lt;/strong&gt;. Kubemoot
exposes tools through the &lt;a href="https://modelcontextprotocol.io"&gt;Model Context Protocol (MCP)&lt;/a&gt;,
so any MCP server becomes a capability a crew can use.&lt;/p&gt;
&lt;h2 id="servers-gateway-agents"&gt;Servers, gateway, agents&lt;/h2&gt;
&lt;p&gt;Three resources form the tool chain, and agents never connect to a tool server
directly:&lt;/p&gt;
&lt;pre tabindex="0"&gt;&lt;code&gt;Agent → MCPGateway → MCPServer
&lt;/code&gt;&lt;/pre&gt;&lt;ul&gt;
&lt;li&gt;An &lt;strong&gt;MCPServer&lt;/strong&gt; deploys one MCP tool server (a container image) and manages its
lifecycle: the deployment, health probes, and - for the many servers that only
speak stdio - an auto-injected &lt;code&gt;mcp-bridge&lt;/code&gt; sidecar that bridges HTTP/SSE to the
server&amp;rsquo;s stdin/stdout.&lt;/li&gt;
&lt;li&gt;An &lt;strong&gt;MCPGateway&lt;/strong&gt; is the hub agents talk to. It discovers MCPServers by label
selector, performs the MCP handshake with each, aggregates all their tools, and
routes an agent&amp;rsquo;s tool call to the right server.&lt;/li&gt;
&lt;li&gt;An &lt;strong&gt;Agent&lt;/strong&gt; reaches the gateway through its runtime and sees the union of tools the
gateway exposes, narrowed by the agent&amp;rsquo;s own &lt;code&gt;enabledTools&lt;/code&gt; / &lt;code&gt;disabledTools&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This indirection means you can add or remove a tool server, or scale the gateway,
without touching agent specs - the gateway rediscovers servers and re-aggregates tools.&lt;/p&gt;</description></item></channel></rss>