Fast multi-cluster Kubernetes TUI built for real fleet scale.
sufsk8s targets the failure mode where existing tools get slow or unusable once you have thousands of pods, many deployments, and multiple kubeconfig contexts open at once. It keeps informer-backed state in memory, renders only the visible table window, and treats multi-cluster as the default model instead of an afterthought.
Working, daily-usable build. Current scope includes:
- startup context picker with persisted selected contexts
- grouped resource catalog with user-managed favourites
- informer-backed built-in views for pods, deployments, services, nodes
- generic discovered-resource and CRD browsing via dynamic client/informer path
- fuzzy, wildcard, exact, and structured per-column filtering
- stacked multi-column sorting
- quick resource finder, namespace picker, context scope picker
- resource details, manifest editing, exec, pod/deployment/node logs, port-forward, scale, restart
- pod/node resource usage in details and tables
- deterministic model coverage plus isolated live k3d/real-PTY acceptance testing
See docs/POC.md for original goals. See docs/ROADMAP.md for next-step improvements and pod-specific needs. See docs/PERFORMANCE-PLAN.md for the focused performance work plan.
k9s and similar tools degrade badly once object counts get large. Typical failures:
- slow startup from repeated list/poll work
- no real multi-cluster unified view
- poor handling of CRDs
- action workflows that are either missing or too brittle
surfsk8s is built around informer/event flow, bounded rendering work, and cluster-aware state from the start.
- Informers, not polling. Shared informer streams deltas. Full list work only when needed.
- Multi-cluster first. Every row carries cluster identity. Unified fleet view is normal.
- Virtual scrolling. Render visible rows only. Large tables stay responsive.
- Generic resource support. Built-ins get typed views where it matters. Everything else still works generically.
- Bounded background work. Lazy dynamic watches, bounded generic-watch count, cached filtered/sorted slices.
- Local-only operation. Reads kubeconfig directly. No server, DB, or agent sidecar.
┌─────────────────────────────────────────┐
│ Bubble Tea App │
│ ┌─────────┐ ┌──────────┐ ┌─────────┐ │
│ │ Views │ │ Filters │ │ Status │ │
│ │ typed + │ │ fuzzy / │ │ scopes, │ │
│ │ generic │ │ column │ │ health │ │
│ └────┬────┘ └────┬─────┘ └────┬────┘ │
│ └────────────┼───────────┘ │
│ │ │
│ ┌──────┴──────┐ │
│ │ State Store │ │
│ │ indexed, │ │
│ │ cross-cluster│ │
│ └──────┬──────┘ │
│ ┌─────────────┼─────────────┐ │
│ ┌───┴────┐ ┌─────┴────┐ ┌─────┴────┐ │
│ │typed │ │generic │ │actions │ │
│ │informers│ │dyn watches│ │kubectl │ │
│ │pods/etc │ │CRDs/GVRs │ │editor │ │
│ └─────────┘ └───────────┘ └──────────┘ │
└────────────────────────────────────────────┘
- Go
- client-go
- Bubble Tea
- Lip Gloss
surfsk8s/
├── main.go
├── cmd/ # CLI entrypoint, flags
├── internal/
│ ├── app/ # Bubble Tea model, screens, workflows
│ ├── actions/ # kubectl-backed exec/edit/port-forward/scale/restart
│ ├── cluster/ # discovery, typed clients, dynamic clients, watches
│ ├── informer/ # typed informer wrappers
│ ├── state/ # cross-cluster store, indexes, detail snapshots
│ └── ui/
│ ├── components/ # table, filter, statusbar
│ ├── theme/ # colors/styles
│ └── views/ # typed resource views
└── docs/
├── POC.md
├── ROADMAP.md
└── PERFORMANCE-PLAN.md
go run .Flags:
go run . -kubeconfig ~/.kube/config -context aws-dev-virginia -namespace defaultAvailable flags:
-kubeconfigpath to kubeconfig-contextinitial kubeconfig context to load-namespaceinitial namespace scope
Common list keys:
j/k, arrows: moveg/G: top/bottomh/l,H/L: horizontal move / edgeenter: open selected itemesc: clear active prompt or go backq: quit
/global list filter- fuzzy by default
- wildcard with
*/? - exact-substring with
=value
fadd structured column filterFmanage column filtersoadd sortOmanage stacked sortsnnamespace pickerccontext scope pickerareset namespace to allrquick resource finderRquick resource finder from resource details
ycopy selected rowYcopy full filtered table as CSV with headers
Pod details:
xexec shellllogspport-forwardeedit manifest
Service details:
pport-forwardeedit manifest
Deployment details:
llogssscalerrestarteedit manifest
Node details:
llogseedit manifest
Resource list row actions:
- deployments:
Sscale selected row,Rrestart selected row - services:
Pport-forward selected row
Mutating actions use confirmation screens. Multi-container exec/log workflows, multi-port port-forward, and node log file selection use explicit picker flows.
Typed first-class views:
- Pods
- Deployments
- Services
- Nodes
Generic browsable support:
- native discovered resources not covered by typed views
- CRDs via dynamic client path
- CRD
additionalPrinterColumns - generic details and manifest editing
- CPU: live usage from
metrics.k8s.io - Memory: live usage from
metrics.k8s.io - Ephemeral: kubelet summary when reachable
- GPU: allocated GPU count from pod spec requests
Table cell format:
<bar> <used>/<total>- pod total = limit, fallback request
- vertical marker inside bar = request when request < limit
- CPU: live usage / allocatable
- Memory: live usage / allocatable
- Ephemeral: kubelet summary / allocatable when reachable
- GPU: allocated on scheduled pods / allocatable
Table cell format:
<bar> <used>/<total>- node total = allocatable
Missing optional sources stay blank, by design.
Typical cases:
- CPU/memory blank → metrics API missing/unavailable
- ephemeral blank → kubelet summary unavailable or blocked
- GPU blank → no allocatable/allocated GPU resource present
Required:
- valid kubeconfig
- network access to selected clusters
Needed for action workflows:
kubectlinPATH$VISUALor$EDITORfor manifest edit, elsevi
Needed for richer metrics:
metrics.k8s.iofor pod/node CPU and memory- kubelet node proxy summary access for ephemeral usage
- advertised GPU resources on pods/nodes for GPU allocation display
Fast deterministic layer:
go test ./...Live Kubernetes and real-PTY acceptance layer:
./scripts/e2e/run.shThe model suite stays offline and fast. The live suite creates an isolated, uniquely named k3d cluster and never reads or changes the default kubeconfig. Docker access has host-level security impact. Read docs/live-e2e.md for prerequisites, blast-radius details, failure artifacts, CI behavior, and the separate bounded agentic protocol.
- typed views exist for high-value built-ins only; many resources use generic view
- ephemeral usage depends on kubelet summary reachability
- actions shell out through
kubectl, not direct API exec/port-forward streams yet - list usage bars are intentionally compact; exact formatting may still evolve
- live E2E requires a healthy local Docker daemon and supported Linux architecture