<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Architecture on Kubemoot</title><link>https://kubemoot.org/docs/architecture/</link><description>Recent content in Architecture on Kubemoot</description><generator>Hugo</generator><language>en</language><atom:link href="https://kubemoot.org/docs/architecture/index.xml" rel="self" type="application/rss+xml"/><item><title>Kubemoot Architecture - Distributed Consensus vs Single-Model Orchestration</title><link>https://kubemoot.org/docs/architecture/consensus-vs-single-model/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://kubemoot.org/docs/architecture/consensus-vs-single-model/</guid><description>&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;&amp;ldquo;Kubemoot composes capability horizontally - many small specialists, because no single model is big enough to do it all; intelligence is emergent from a committee. A single-model agent is the opposite: capability concentrated in one large reasoner that drives tools and, when needed, spawns copies of itself.&amp;rdquo;&lt;/em&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;a single-model coding agent, contrasting itself with Kubemoot&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;
&lt;h2 id="purpose"&gt;Purpose&lt;/h2&gt;
&lt;p&gt;There are two broad shapes an agentic system can take. Kubemoot is one of them; most well-known agent systems are the other. The fastest way to understand &lt;em&gt;why Kubemoot is built the way it is&lt;/em&gt; - many small models, a coordinator, consensus signals over a message bus - is to hold it next to its opposite and see what each shape buys and costs.&lt;/p&gt;</description></item><item><title>Agentic Consensus</title><link>https://kubemoot.org/docs/architecture/agentic-consensus/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://kubemoot.org/docs/architecture/agentic-consensus/</guid><description>&lt;p&gt;Kubemoot answers a question by convening the agents whose expertise fits it, letting
them investigate and deliberate, and settling on a synthesis once the crew has spoken.
This page is the architecture behind that: how the coordinator selects who takes part,
what a consensus signal means to the protocol, and why a discussion settles by state
rather than by a clock. For the concept-level model (roles, phases, signal vocabulary)
see &lt;a href="../../concepts/crews-and-agents/"&gt;Crews &amp;amp; Agents&lt;/a&gt;,
&lt;a href="../../concepts/consensus-model/"&gt;The Moot&lt;/a&gt;, and
&lt;a href="../../concepts/signals-and-protocol/"&gt;Signals &amp;amp; Protocol&lt;/a&gt; - this page assumes them and
goes one level deeper into the coordinator&amp;rsquo;s mechanics.&lt;/p&gt;</description></item><item><title>Tooler-Analyst Architecture</title><link>https://kubemoot.org/docs/architecture/tooler-analyst-architecture/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://kubemoot.org/docs/architecture/tooler-analyst-architecture/</guid><description>&lt;h2 id="two-roles-that-serve-different-phases"&gt;Two Roles That Serve Different Phases&lt;/h2&gt;
&lt;p&gt;Kubemoot organizes the data-gathering and reasoning work of a discussion into two complementary roles that mirror how real teams operate: &lt;strong&gt;Toolers&lt;/strong&gt; who interact with live systems, and &lt;strong&gt;Analysts&lt;/strong&gt; who reason over what the Toolers found.&lt;/p&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;&lt;/th&gt;
 &lt;th&gt;&lt;strong&gt;Toolers (MCP)&lt;/strong&gt;&lt;/th&gt;
 &lt;th&gt;&lt;strong&gt;Analysts (RAG)&lt;/strong&gt;&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;&lt;strong&gt;Phase&lt;/strong&gt;&lt;/td&gt;
 &lt;td&gt;EVALUATING&lt;/td&gt;
 &lt;td&gt;REVIEW&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;&lt;strong&gt;Declared&lt;/strong&gt;&lt;/td&gt;
 &lt;td&gt;Agents with MCP tools&lt;/td&gt;
 &lt;td&gt;Agents with embedded domain docs&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;&lt;strong&gt;Self-discovered&lt;/strong&gt;&lt;/td&gt;
 &lt;td&gt;Onboarding agent deploys new MCP servers&lt;/td&gt;
 &lt;td&gt;RTFM agent indexes new documentation&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;&lt;strong&gt;Carries&lt;/strong&gt;&lt;/td&gt;
 &lt;td&gt;MCP tools, no live tools in reasoning&lt;/td&gt;
 &lt;td&gt;RAG sources, no MCP tools&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;&lt;strong&gt;Thinking&lt;/strong&gt;&lt;/td&gt;
 &lt;td&gt;OFF (deterministic tool selection)&lt;/td&gt;
 &lt;td&gt;ON (reasoning over gathered data)&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;&lt;strong&gt;Analogy&lt;/strong&gt;&lt;/td&gt;
 &lt;td&gt;The person with their hands on the keyboard&lt;/td&gt;
 &lt;td&gt;The person with the manual open, weighing the evidence&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;&lt;strong&gt;GPU profile&lt;/strong&gt;&lt;/td&gt;
 &lt;td&gt;Heavier (tool-calling inference)&lt;/td&gt;
 &lt;td&gt;Lighter (reasoning-focused, no tool overhead)&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;h3 id="why-separate-them"&gt;Why Separate Them&lt;/h3&gt;
&lt;p&gt;LLMs have a practical limit on how much context they can process while still reliably selecting tools and interpreting results. When a single agent carries both MCP tool schemas and RAG documentation in its system prompt, the prompt bloats and tool selection degrades. Separating concerns gives each agent a smaller, focused prompt: Toolers are better at picking the right tool, Analysts are better at weighing the evidence from those tools.&lt;/p&gt;</description></item><item><title>Scheduling - reminders, follow-ups, recurring</title><link>https://kubemoot.org/docs/architecture/scheduling/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://kubemoot.org/docs/architecture/scheduling/</guid><description>&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;One sentence&lt;/strong&gt;: a crew&amp;rsquo;s &lt;code&gt;scheduler-advisor&lt;/code&gt; Tooler agent calls the &lt;code&gt;scheduling-mcp&lt;/code&gt; server&amp;rsquo;s tools to write &lt;code&gt;ScheduleRecord&lt;/code&gt;s into the NATS KV bucket &lt;code&gt;kubemoot_scheduled&lt;/code&gt;; the operator&amp;rsquo;s scheduler poller fires due records - either inline in the originating thread or as a new thread, depending on &lt;code&gt;kind&lt;/code&gt; x whether a &lt;code&gt;sourceThreadId&lt;/code&gt; was attached.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="why-this-exists"&gt;Why this exists&lt;/h2&gt;
&lt;p&gt;Some discussions are inherently &lt;em&gt;temporally&lt;/em&gt; driven:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Reminder&lt;/strong&gt; - &amp;ldquo;Remind me in 30 minutes to check the backup job status.&amp;rdquo;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Follow-up&lt;/strong&gt; - &amp;ldquo;Re-check disk usage in an hour&amp;rdquo; / &amp;ldquo;Run the weekly health summary every Monday.&amp;rdquo;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Recurring&lt;/strong&gt; - daily/weekly admin checks, configured by a cluster admin rather than via chat.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;All three reduce to: &lt;em&gt;publish a discussion message at a future time, in the right place&lt;/em&gt;. This module provides the primitive.&lt;/p&gt;</description></item><item><title>Kubemoot Scheduler</title><link>https://kubemoot.org/docs/architecture/scheduler/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://kubemoot.org/docs/architecture/scheduler/</guid><description>&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Audience:&lt;/strong&gt; operator and agent-runtime developers; developers of the ecosystem tools (CrewForge, kmctl); open-source consumers.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="one-paragraph-summary"&gt;One-Paragraph Summary&lt;/h2&gt;
&lt;p&gt;Kubemoot is an LLM orchestrator. Like &lt;code&gt;kube-scheduler&lt;/code&gt; matches Pods to Nodes based on resource requests, the &lt;strong&gt;Kubemoot scheduler matches agents to &lt;code&gt;(Model, Provider)&lt;/code&gt; bindings based on capability requirements.&lt;/strong&gt; Agents declare what they can do (capabilities), Models declare what they are (family, params, capabilities) as labels, ModelProviders declare what they have (VRAM, parallelism), and a &lt;code&gt;CrewSchedulingPolicy&lt;/code&gt; CRD expresses require/prefer rules using Kubernetes-style label selectors. The match is computed in the agent reconciler via filter → score → bind; the resolved &lt;code&gt;(model, provider, endpoint)&lt;/code&gt; is applied to the agent&amp;rsquo;s Deployment.&lt;/p&gt;</description></item></channel></rss>