Skip to content

feat(monitor): Prometheus + Grafana monitoring stack (GPU, node, cluster) - #28

Closed
MustafaOgur wants to merge 1 commit into
mainfrom
devops/feature
Closed

MustafaOgur wants to merge 1 commit into
mainfrom
devops/feature

Conversation

@MustafaOgur

Copy link
Copy Markdown
Collaborator

Summary

Self-contained monitoring stack in deephorizon-monitor, deployed via GitOps:
Prometheus + node-exporter + kube-state-metrics + Grafana, with a cluster dashboard
as code. Monitors GPU (dcgm), node, cluster objects, and any pod annotated
prometheus.io/scrape: "true" (api/inference auto-join later).

Details, access, and the post-merge step → docs/MONITORING.md.

Testing

Verified end-to-end on the cluster — all scrape targets up (incl. GPU), Grafana and
the provisioned dashboard live. Removed afterwards; this PR is the clean GitOps deploy.

Notes

  • Required after merge: apply the grafana-admin SealedSecret (password not in Git) — see docs.
  • Images pinned (no latest), non-root pods, read-only RBAC.
  • Grafana NodePort 30030 needs a UFW LAN rule; GPU/Node dashboards (12239 / 1860) imported manually.

@Enskc05 Enskc05 closed this Aug 3, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants