<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Operating on Kubemoot</title><link>https://kubemoot.org/docs/operating/</link><description>Recent content in Operating on Kubemoot</description><generator>Hugo</generator><language>en</language><atom:link href="https://kubemoot.org/docs/operating/index.xml" rel="self" type="application/rss+xml"/><item><title>Onboarding Guide</title><link>https://kubemoot.org/docs/operating/onboarding-guide/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://kubemoot.org/docs/operating/onboarding-guide/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;When no agent in a crew can answer a question, the coordinator signals a capability gap. An agent running in onboarding mode reacts to that signal: it proposes an MCP server that could fill the gap. On the user&amp;rsquo;s consent it replies with a proposed &lt;code&gt;MCPServer&lt;/code&gt; manifest in the discussion thread. Applying that manifest, and creating the Tooler &lt;code&gt;Agent&lt;/code&gt; that uses it, are manual steps today.&lt;/p&gt;
&lt;p&gt;This page describes what ships today, how to enable it, and what is not automated yet. For the consensus signals that produce a gap, see &lt;a href="../../architecture/agentic-consensus/"&gt;Agentic Consensus&lt;/a&gt;.&lt;/p&gt;</description></item><item><title>Kubemoot Dashboard</title><link>https://kubemoot.org/docs/operating/dashboard/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://kubemoot.org/docs/operating/dashboard/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;The Kubemoot Dashboard is a web UI for looking at a running Kubemoot installation. It lists Kubemoot resources across namespaces, shows GPU node and model state, and streams agent discussions and NATS messages as they happen. It is a SvelteKit application that reads the Kubernetes API and the NATS bus from the server side; the browser never connects to either directly.&lt;/p&gt;
&lt;p&gt;Metrics, logs, and traces are not part of the dashboard. For those, see &lt;a href="../observability/"&gt;Observability&lt;/a&gt;.&lt;/p&gt;</description></item><item><title>Kubemoot Observability</title><link>https://kubemoot.org/docs/operating/observability/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://kubemoot.org/docs/operating/observability/</guid><description>&lt;p&gt;How to observe what a Kubemoot crew is doing: the data it exposes, the surfaces that render it, and the external observability stack it depends on. For the dashboard UI itself see &lt;a href="https://kubemoot.org/docs/operating/dashboard/"&gt;dashboard.md&lt;/a&gt;; for the consensus protocol that produces most of this data see &lt;a href="https://kubemoot.org/docs/architecture/agentic-consensus/"&gt;agentic-consensus.md&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id="what-you-can-observe"&gt;What you can observe&lt;/h2&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;Domain&lt;/th&gt;
 &lt;th&gt;Signals&lt;/th&gt;
 &lt;th&gt;Where it comes from&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;&lt;strong&gt;Discussions / consensus&lt;/strong&gt;&lt;/td&gt;
 &lt;td&gt;Thread lifecycle, per-agent signals (&lt;code&gt;triaging&lt;/code&gt;, &lt;code&gt;evaluating&lt;/code&gt;, &lt;code&gt;waiting&lt;/code&gt;, &lt;code&gt;agree&lt;/code&gt;, &lt;code&gt;concern&lt;/code&gt;, &lt;code&gt;stand_aside&lt;/code&gt;, &lt;code&gt;failure&lt;/code&gt;, &lt;code&gt;block&lt;/code&gt;, &lt;code&gt;advisory&lt;/code&gt;, &lt;code&gt;synthesis&lt;/code&gt;), settle decisions, synthesis text&lt;/td&gt;
 &lt;td&gt;Coordinator + agents publish to NATS &lt;code&gt;kubemoot.discuss.&amp;lt;namespace&amp;gt;.&amp;lt;crew&amp;gt;.*&lt;/code&gt;&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;&lt;strong&gt;Agent activity&lt;/strong&gt;&lt;/td&gt;
 &lt;td&gt;Heartbeats while GPU-queued, triage vs evaluation phases, per-call provider attribution (which model/provider/GPU served each call)&lt;/td&gt;
 &lt;td&gt;Agent runtime signals + metadata&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;&lt;strong&gt;GPU &amp;amp; scheduling&lt;/strong&gt;&lt;/td&gt;
 &lt;td&gt;Loaded-model footprints, in-flight tickets, provider readiness, total/used VRAM, scheduling decisions&lt;/td&gt;
 &lt;td&gt;Operator-published provider state (NATS KV) + DCGM&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;&lt;strong&gt;Quality&lt;/strong&gt;&lt;/td&gt;
 &lt;td&gt;MCP tool quality verdicts (hot/failing/out-of-norm)&lt;/td&gt;
 &lt;td&gt;MCP quality pipeline → &lt;code&gt;kubemoot.quality.*&lt;/code&gt;&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;&lt;strong&gt;RAG indexing&lt;/strong&gt;&lt;/td&gt;
 &lt;td&gt;Source indexing progress, chunk/embed status&lt;/td&gt;
 &lt;td&gt;RAGSource controller&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;&lt;strong&gt;Fitness&lt;/strong&gt;&lt;/td&gt;
 &lt;td&gt;Per-scenario/iteration pass-rates, assertion results, durations, full conversation transcripts&lt;/td&gt;
 &lt;td&gt;&lt;code&gt;CrewFitness&lt;/code&gt; / &lt;code&gt;CrewFitnessSuite&lt;/code&gt; → XLSX in NATS Object Store&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 id="surfaces"&gt;Surfaces&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Kubemoot Dashboard&lt;/strong&gt;: the primary real-time view (discussion threads &amp;amp; timeline, agent topology, NATS messages, fitness suites, GPU/model activity). Updates push-style via NATS JetStream SSE and k8s watch→SSE. See &lt;a href="https://kubemoot.org/docs/operating/dashboard/"&gt;dashboard.md&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;NATS JetStream&lt;/strong&gt;: the event substrate. Streams: &lt;code&gt;KUBEMOOT_CHRONICLE&lt;/code&gt;, &lt;code&gt;KUBEMOOT_QUALITY&lt;/code&gt;, &lt;code&gt;KUBEMOOT_CHAT&lt;/code&gt;, &lt;code&gt;KUBEMOOT_OPERATOR&lt;/code&gt;, &lt;code&gt;KUBEMOOT_DISCUSS&lt;/code&gt;. Anything the dashboard shows is reconstructable from these.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Prometheus metrics&lt;/strong&gt;: the agent runtime exposes Micrometer/Prometheus-format metrics; &lt;code&gt;DiscussionMetrics&lt;/code&gt; (&lt;code&gt;ai.kubemoot.agent.nats&lt;/code&gt;) covers inference timings, token usage, signal counts, and thread/synthesis completion.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="the-three-pillars-of-observability"&gt;The three pillars of observability&lt;/h2&gt;
&lt;p&gt;Kubemoot leans on the cluster&amp;rsquo;s existing observability stack rather than shipping its own. The three pillars (&lt;strong&gt;metrics, logs, spans&lt;/strong&gt;) are at different stages.&lt;/p&gt;</description></item></channel></rss>