kubeagent
v0.1.62
Published
AI-powered Kubernetes management CLI
Readme
KubeAgent
AI-powered Kubernetes monitoring and auto-remediation CLI. Detects cluster issues, investigates root causes, and applies fixes — with human approval for risky actions.
Built for solo DevOps engineers and small teams who want an intelligent on-call assistant.
Quick Start
# Install
npm install -g kubeagent
# Log in (opens browser; uses device-code flow on headless machines)
kubeagent login
# Onboard — scan cluster, detect your services, build knowledge base
kubeagent onboard
# Watch mode — continuous monitoring with auto-remediation
kubeagent watchUpdating
Upgrade the CLI in place with the package manager that installed it (npm, pnpm, bun, or yarn — detected automatically):
kubeagent update # upgrade to the latest version
kubeagent update --check # report only; exit 0 up to date, 1 update available, 2 registry unreachableCommands
| Command | Description |
|---------|-------------|
| status | Quick cluster health check (no LLM) |
| onboard | Scan cluster + codebases, generate knowledge base |
| watch | Continuous monitoring with auto-remediation |
| agent install / uninstall / status | Install, remove, or check the always-on in-cluster agent (Deployment) |
| check | One-shot cluster health check for CI pipelines (read-only by default, JSON output, exit codes) |
| demo | Safe demo incident: breaks a pod in an isolated kubeagent-demo namespace, watches KubeAgent detect, alert, and diagnose it, then cleans up |
| diagnose <resource> | One-shot diagnosis of a pod/deployment/service |
| scan <directory> | Match local project directories to cluster deployments |
| scan-apps | Detect OSS apps in the cluster, list affected workloads + ingress domains, check for CVEs / EOL (--report for AI-prioritized remediation) |
| notify | Manage notification channels (list, add, remove, test) |
| login | Log in to KubeAgent (browser or --device for headless) |
| account | Show account info and token balance |
| logout | Clear saved credentials |
| update | Upgrade the CLI to the latest version via the package manager that installed it (--check reports only) |
| <prompt> | Ask a freeform question about your cluster |
Common Options
-c, --context <name>— Kubernetes context (defaults to current)-n, --namespace <name>— Namespace (fordiagnose)-i, --interval <seconds>— Check interval forwatch(default: 300)--no-interactive— Auto-deny all approvals, skip questions (forwatchin background/CI)
In-Cluster Agent
kubeagent watch needs a laptop to stay awake. For always-on monitoring, install the agent as a Deployment:
kubeagent login # if not already logged in
kubeagent agent install # mints a dedicated API key and deploys the agent
kubeagent agent status # pod health + recent logs
kubeagent agent uninstall # remove itinstall resolves the cluster name from --cluster-name or the current kubectl context, prints a plan, and — unless --yes skips the prompt — asks for confirmation. It then mints a dedicated API key scoped to that cluster (agent:<cluster-name>), creates the namespace, Secret, RBAC, and Deployment, and waits for the rollout. A reinstall mints a new key and only revokes the previous one after the new rollout succeeds — so a failed reinstall never leaves the cluster without a working key.
uninstall removes the Deployment, Secret, ServiceAccount, and the cluster-scoped ClusterRole and ClusterRoleBinding, then revokes the agent:<cluster-name> API key. Pass --delete-namespace to also remove the namespace.
| Flag | Default | Effect |
|------|---------|--------|
| -c, --context <name> | current context | Kubernetes context to target |
| -n, --namespace <name> | kubeagent | Namespace to install into |
| --cluster-name <name> | the kubectl context in use | Display name in alerts and the API key name |
| --interval <seconds> | 300 | Check interval (minimum 60) |
| --no-auto-fix | auto-fix on | Read-only monitoring; safe remediations are disabled |
| --read-only | off | Grants a ClusterRole with no delete/patch verbs |
| --image-tag <tag> | latest | Pin a specific agent image tag |
| -y, --yes | off | Skip the confirmation prompt |
Prefer Helm or a single manifest instead? Both remain available:
# Helm
helm install kubeagent ./charts/kubeagent-agent \
--namespace kubeagent --create-namespace \
--set apiKey=<API key from ~/.kubeagent/auth.json after `kubeagent login`>
# One-file manifest
kubectl create namespace kubeagent
kubectl -n kubeagent create secret generic kubeagent-agent --from-literal=api-key=<your API key>
kubectl apply -f https://kubeagent.net/install/kubeagent-agent.yamlThe CLI, Helm chart, and one-file manifest deploy the same image and RBAC — pick whichever fits your workflow.
CI Usage
kubeagent check is the one-shot mode for CI pipelines. It scans the cluster, prints a report, and exits with a code your build can gate on — no daemon, no prompts, no onboarding required.
# After deploying: confirm pods reached Ready
kubectl rollout status deployment/my-app -n my-ns --timeout=5m
# Then confirm the system is actually healthy
kubeagent check -n my-nsThe two steps are complementary: rollout status answers "did Kubernetes accept the deploy?", kubeagent check answers "is the system actually healthy?" — catching crashloops, evicted pods, failed jobs, and excessive ephemeral-storage usage that can evict a database even while its pod is still Ready.
Ephemeral-storage checks read kubelet node.fs and /stats/summary data to report low node root-filesystem headroom, per-container writable-layer/log usage, and local-volume or other usage not attributable to a container. CLI checks use the current kubectl identity and fail open when stats are unavailable. In-cluster collection is off by default because Kubernetes does not offer stats-only API-server RBAC: get nodes/proxy can reach other kubelet GET endpoints and may bypass normal admission controls. Trusted operators can explicitly opt in with Helm --set ephemeralStorage.enabled=true; the one-file install remains disabled.
Exit codes: 0 = no issues at/above threshold · 1 = issues found · 2 = kubeagent error (unreachable cluster, bad flag, timeout).
Common flags:
-n, --namespace <ns>— scope to one or more namespaces (repeatable)--fail-on <severity>— fail-threshold:info,warning, orcritical(default:critical)--format json— emit a versioned JSON report instead of human output--diagnose— run LLM root-cause analysis on findings (requiresKUBEAGENT_API_KEYenv var)
See docs/CI.md for full reference, GitHub Actions and GitLab CI snippets, and the JSON schema.
Knowledge Base
KubeAgent builds a per-cluster knowledge base at ~/.kubeagent/clusters/<context>/ during onboarding:
~/.kubeagent/
config.yaml # Global config (clusters, webhooks, intervals)
clusters/
hetzner-prod/
cluster.md # Nodes, namespaces, deployments, services
applications.md # Detected OSS apps with CVE / EOL status
projects/
api.md # Tech stack, dependencies, notes
runbooks/ # Custom runbooks (manually added)
incidents/ # Auto-logged incident reportsAction Safety
Actions are classified into tiers:
| Tier | Actions | Behavior | |------|---------|----------| | Safe | get logs, describe, get events, restart pod, rollout restart, scale up | Auto-executed | | Risky | rollback, scale to zero, delete resource, edit configmap, apply manifest | Requires approval | | Never | delete namespace, delete node, etc. | Refused |
Notifications
Add and manage channels from the CLI:
kubeagent notify add # interactive setup
kubeagent notify list # show configured channels
kubeagent notify test # send a test alert to all channels
kubeagent notify remove 1 # remove channel by indexSupported channels:
| Channel | Setup | |---------|-------| | Slack | Connect via OAuth from the dashboard at app.kubeagent.net, or paste a webhook URL | | PagerDuty | Add your integration key — verified automatically | | Discord | Paste a Discord webhook URL | | Microsoft Teams | Paste a Teams incoming webhook URL | | Custom Webhook | Any HTTP endpoint that accepts JSON payloads |
Observability
The SaaS server exposes:
GET /healthfor Kubernetes readiness/liveness probesGET /metricsfor Prometheus scraping
Kubernetes monitoring manifests live in k8s/servicemonitor.yaml and k8s/prometheusrule.yaml. Internal rollout notes are in docs/MONITORING.md.
Requirements
- Node.js >= 22
kubectlconfigured with cluster access- A KubeAgent account — sign up at app.kubeagent.net
FAQ
Can I monitor multiple clusters?
Yes. Run a separate kubeagent watch instance for each cluster, specifying the context with -c:
kubeagent watch -c hetzner-prod
kubeagent watch -c staging-clusterEach instance operates independently and reports incidents to the same account.
Does kubeagent modify my cluster without asking? Safe actions (pod restarts, rollout restarts, scaling up) are applied automatically. Anything potentially destructive pauses and waits for your approval.
Where is my API key stored?
Credentials are stored locally at ~/.kubeagent/auth.json and are never sent anywhere except the KubeAgent API.
Links
- Homepage
- Blog — incident walkthroughs, agentic-diagnosis patterns, on-call economics (RSS)
- Docs
- Sign up
License
Proprietary — see LICENSE for details. Commercial use requires a paid plan at kubeagent.net.
