Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
85 changes: 85 additions & 0 deletions docs/MONITORING.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,85 @@
# Monitoring — Prometheus + Grafana

Cluster, node ve **GPU** metriklerini toplar (Prometheus), gösterir (Grafana).
Kurulum GitOps ile: `infra/k8s/monitor/`, `apps/monitor.yaml`, namespace
`deephorizon-monitor`.

## Erişim

| | |
|:---|:---|
| **Grafana (LAN)** | `http://10.10.1.132:30030` |
| **Kullanıcı** | `admin` — parola DevOps'ta (parola kasası) |
| **Prometheus** | ClusterIP (dışarı açılmaz); Grafana'ya datasource olarak önceden bağlı |

UFW 30030'u yalnızca yerel subnet'e açar. İleride NPM arkasına domain'le alınabilir.

## Ne izler

| Kaynak | Ne verir |
|:---|:---|
| **node-exporter** | node CPU / RAM / disk / ağ |
| **kube-state-metrics** | pod/deployment/PVC/job durumları (restart, hazır replika…) |
| **cAdvisor** (kubelet) | konteyner kaynak kullanımı |
| **dcgm-exporter** (GPU operator) | L40S kullanım / VRAM / sıcaklık / güç |
| **anotasyon keşfi** | `prometheus.io/scrape: "true"` taşıyan her pod (api, inference) |

> dcgm-exporter anotasyon taşımadığı için özel bir scrape job'ı ile toplanır
> (`gpu-operator-resources/nvidia-dcgm-exporter`). api/inference deploy olunca
> anotasyonla otomatik girer — ekstra iş yok.

## Dashboard'lar

- **DeepHorizon - Cluster** — koddan otomatik gelir (`infra/k8s/monitor/dashboards/deephorizon-cluster.json`,
Grafana provisioning). Pod/CPU/RAM/restart + deployment durumu, bizim etiket şemamıza göre.
- **GPU + Node** — topluluk panoları, repoya gömülmedi (JSON büyük). Deploy sonrası
**elle import** (Grafana → New → Import → ID → datasource `Prometheus`):
- **12239** — NVIDIA DCGM (GPU: kullanım, VRAM, sıcaklık, güç)
- **1860** — Node Exporter Full (CPU / RAM / disk / ağ)

## Grafana admin parolası — elle (Git'e girmez)

Grafana pod'u `grafana-admin` Secret'ı olmadan başlamaz. Secret'lar Git dışında;
SealedSecret ile bir kez uygulanır:

```bash
# 1. Duz secret taslagi (commit etme)
kubectl create secret generic grafana-admin \
--namespace deephorizon-monitor \
--from-literal=admin-password='<guclu-parola>' \
--dry-run=client -o yaml > /tmp/grafana-admin.yaml

# 2. Muhurle (namespace scope zorunlu)
kubeseal -n deephorizon-monitor -o yaml \
< /tmp/grafana-admin.yaml > /tmp/grafana-admin-sealed.yaml

# 3. Cluster'a uygula (Git'e DEGIL)
kubectl apply -f /tmp/grafana-admin-sealed.yaml

# 4. Temizle; muhurlu YAML'i parola kasasina koy
shred -u /tmp/grafana-admin.yaml /tmp/grafana-admin-sealed.yaml
```

> **Sync sırası:** Grafana bu secret'a bağlı. Argo CD namespace'i oluşturduktan
> sonra secret'ı apply et; gelene kadar Grafana `CreateContainerConfigError`'da
> bekler, secret gelince kendi düzelir (takılırsa `kubectl delete pod -n
> deephorizon-monitor -l app=grafana`).

## DevOps notları

| | |
|:---|:---|
| Manifest'ler | `infra/k8s/monitor/` (kustomize) |
| Argo CD app | `infra/k8s/apps/monitor.yaml` |
| Grafana NodePort | `30030` (UFW: `ufw allow from 10.10.1.0/24 to any port 30030 proto tcp`) |
| Secret | `grafana-admin` (SealedSecret, Git dışı) |
| Dashboard'lar | `infra/k8s/monitor/dashboards/*.json` → `grafana-dashboards` ConfigMap (generator) |

**Yeni scrape hedefi eklemek:** pod'a `prometheus.io/scrape: "true"` +
`prometheus.io/port: "<port>"` anotasyonu koy → otomatik toplanır. Anotasyon
konamıyorsa (üçüncü parti) `prometheus.yaml`'e özel bir job ekle (dcgm örneği gibi),
yeniden render/apply et.

**Yeni dashboard eklemek:** JSON'ı `infra/k8s/monitor/dashboards/`'a koy →
`kustomization.yaml`'deki `configMapGenerator.files` listesine ekle → Grafana
provider onu otomatik yükler.
1 change: 1 addition & 0 deletions docs/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -11,6 +11,7 @@ Project documentation that does not belong in source-level READMEs.
| [`DVC.md`](DVC.md) | **DVC rehberi (TR)** — veri versiyonlama: `dvc-cache` remote'unun kurulumu (`dvc init` + `remote add`), push/pull kullanımı, eski sürüme dönme. **DVC kullanacak kişinin adresi.** |
| [`AIRFLOW.md`](AIRFLOW.md) | **Airflow rehberi (TR)** — veri pipeline orkestrasyonu: erişim, DAG teslimi (git-sync), çalışma ortamı kontratı (env, workspace, imaj), Data squad'dan kalanlar. **DAG yazacak kişinin adresi.** |
| [`DEVOPS.md`](DEVOPS.md) | **DevOps kurulum günlüğü (TR)** — GPU sunucu bootstrap'ının tamamı: NVIDIA sürücü, MicroK8s + GPU addon, Sealed Secrets, Argo CD, NPM, UFW; karşılaşılan hatalar, sebepleri ve çözümleri. |
| [`MONITORING.md`](MONITORING.md) | **Monitoring rehberi (TR)** — Prometheus + Grafana + exporter'lar + GPU (dcgm): erişim, ne izlenir, dashboard'lar, `grafana-admin` SealedSecret akışı, yeni scrape/dashboard ekleme. |
| `adr/` | Architecture Decision Records — one file per non-trivial decision |
| `runbooks/` | On-call / incident playbooks (one per failure mode) |

Expand Down
22 changes: 22 additions & 0 deletions infra/k8s/apps/monitor.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,22 @@
# Monitoring stack (kustomize). Grafana admin parolasi Git'te degil:
# 'grafana-admin' SealedSecret'i elle apply edilmeli (bkz. docs/MONITORING.md).
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: monitor
namespace: argocd
spec:
project: default
source:
repoURL: https://github.com/Octapull/deephorizon.git
targetRevision: main
path: infra/k8s/monitor
destination:
server: https://kubernetes.default.svc
namespace: deephorizon-monitor
syncPolicy:
automated:
prune: true
selfHeal: true
syncOptions:
- CreateNamespace=true
133 changes: 133 additions & 0 deletions infra/k8s/monitor/dashboards/deephorizon-cluster.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,133 @@
{
"title": "DeepHorizon - Cluster",
"uid": "deephorizon-cluster",
"schemaVersion": 39,
"version": 1,
"editable": true,
"time": { "from": "now-6h", "to": "now" },
"refresh": "30s",
"templating": {
"list": [
{
"name": "datasource",
"type": "datasource",
"query": "prometheus",
"current": {},
"hide": 0
},
{
"name": "namespace",
"type": "query",
"datasource": { "type": "prometheus", "uid": "${datasource}" },
"query": "label_values(kube_pod_info, namespace)",
"definition": "label_values(kube_pod_info, namespace)",
"refresh": 2,
"includeAll": true,
"multi": true,
"current": { "text": "All", "value": "$__all" }
}
]
},
"panels": [
{
"id": 1,
"type": "stat",
"title": "Running Pods",
"gridPos": { "h": 4, "w": 6, "x": 0, "y": 0 },
"datasource": { "type": "prometheus", "uid": "${datasource}" },
"options": { "colorMode": "value", "graphMode": "area", "reduceOptions": { "calcs": ["lastNotNull"] } },
"targets": [
{ "refId": "A", "datasource": { "type": "prometheus", "uid": "${datasource}" },
"expr": "sum(kube_pod_status_phase{phase=\"Running\"})" }
]
},
{
"id": 2,
"type": "stat",
"title": "Pods Not Running",
"gridPos": { "h": 4, "w": 6, "x": 6, "y": 0 },
"datasource": { "type": "prometheus", "uid": "${datasource}" },
"fieldConfig": { "defaults": { "thresholds": { "mode": "absolute", "steps": [ { "color": "green", "value": null }, { "color": "red", "value": 1 } ] } }, "overrides": [] },
"options": { "colorMode": "value", "graphMode": "none", "reduceOptions": { "calcs": ["lastNotNull"] } },
"targets": [
{ "refId": "A", "datasource": { "type": "prometheus", "uid": "${datasource}" },
"expr": "sum(kube_pod_status_phase{phase!=\"Running\"})" }
]
},
{
"id": 3,
"type": "stat",
"title": "Container Restarts (total)",
"gridPos": { "h": 4, "w": 6, "x": 12, "y": 0 },
"datasource": { "type": "prometheus", "uid": "${datasource}" },
"options": { "colorMode": "value", "graphMode": "area", "reduceOptions": { "calcs": ["lastNotNull"] } },
"targets": [
{ "refId": "A", "datasource": { "type": "prometheus", "uid": "${datasource}" },
"expr": "sum(kube_pod_container_status_restarts_total)" }
]
},
{
"id": 4,
"type": "stat",
"title": "Nodes Ready",
"gridPos": { "h": 4, "w": 6, "x": 18, "y": 0 },
"datasource": { "type": "prometheus", "uid": "${datasource}" },
"options": { "colorMode": "value", "graphMode": "none", "reduceOptions": { "calcs": ["lastNotNull"] } },
"targets": [
{ "refId": "A", "datasource": { "type": "prometheus", "uid": "${datasource}" },
"expr": "sum(kube_node_status_condition{condition=\"Ready\",status=\"true\"})" }
]
},
{
"id": 5,
"type": "timeseries",
"title": "CPU by pod (cores, top 10)",
"gridPos": { "h": 8, "w": 12, "x": 0, "y": 4 },
"datasource": { "type": "prometheus", "uid": "${datasource}" },
"fieldConfig": { "defaults": { "unit": "short" }, "overrides": [] },
"targets": [
{ "refId": "A", "datasource": { "type": "prometheus", "uid": "${datasource}" },
"legendFormat": "{{namespace}}/{{pod}}",
"expr": "topk(10, sum by (namespace, pod) (rate(container_cpu_usage_seconds_total{container!=\"\", namespace=~\"$namespace\"}[5m])))" }
]
},
{
"id": 6,
"type": "timeseries",
"title": "Memory by pod (working set, top 10)",
"gridPos": { "h": 8, "w": 12, "x": 12, "y": 4 },
"datasource": { "type": "prometheus", "uid": "${datasource}" },
"fieldConfig": { "defaults": { "unit": "bytes" }, "overrides": [] },
"targets": [
{ "refId": "A", "datasource": { "type": "prometheus", "uid": "${datasource}" },
"legendFormat": "{{namespace}}/{{pod}}",
"expr": "topk(10, sum by (namespace, pod) (container_memory_working_set_bytes{container!=\"\", namespace=~\"$namespace\"}))" }
]
},
{
"id": 7,
"type": "table",
"title": "Deployments - available replicas",
"gridPos": { "h": 8, "w": 12, "x": 0, "y": 12 },
"datasource": { "type": "prometheus", "uid": "${datasource}" },
"targets": [
{ "refId": "A", "datasource": { "type": "prometheus", "uid": "${datasource}" },
"format": "table", "instant": true,
"expr": "kube_deployment_status_replicas_available{namespace=~\"$namespace\"}" }
]
},
{
"id": 8,
"type": "timeseries",
"title": "Restart rate by pod (15m)",
"gridPos": { "h": 8, "w": 12, "x": 12, "y": 12 },
"datasource": { "type": "prometheus", "uid": "${datasource}" },
"fieldConfig": { "defaults": { "unit": "short" }, "overrides": [] },
"targets": [
{ "refId": "A", "datasource": { "type": "prometheus", "uid": "${datasource}" },
"legendFormat": "{{namespace}}/{{pod}}",
"expr": "sum by (namespace, pod) (rate(kube_pod_container_status_restarts_total{namespace=~\"$namespace\"}[15m]))" }
]
}
]
}
133 changes: 133 additions & 0 deletions infra/k8s/monitor/grafana.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,133 @@
# Grafana — Prometheus datasource'u onceden tanimli gelir (provisioning).
# Admin parolasi repoda YOK: 'grafana-admin' SealedSecret (bkz. docs/MONITORING.md).
apiVersion: v1
kind: ConfigMap
metadata:
name: grafana-datasources
namespace: deephorizon-monitor
data:
datasource.yaml: |
apiVersion: 1
datasources:
- name: Prometheus
type: prometheus
access: proxy
url: http://prometheus.deephorizon-monitor.svc:9090
isDefault: true
---
# Dashboard provider — /etc/grafana/dashboards'daki JSON'lari otomatik yukler.
apiVersion: v1
kind: ConfigMap
metadata:
name: grafana-dashboard-provider
namespace: deephorizon-monitor
data:
provider.yaml: |
apiVersion: 1
providers:
- name: deephorizon
orgId: 1
type: file
disableDeletion: true
updateIntervalSeconds: 30
allowUiUpdates: true
options:
path: /etc/grafana/dashboards
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: grafana-data
namespace: deephorizon-monitor
spec:
accessModes: ["ReadWriteOnce"]
resources:
requests:
storage: 5Gi
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: grafana
namespace: deephorizon-monitor
spec:
replicas: 1
strategy:
type: Recreate
selector:
matchLabels:
app: grafana
template:
metadata:
labels:
app: grafana
spec:
securityContext:
# Grafana imaji 472 kullanicisiyla kosar; PVC'yi bu gruba yazilabilir yap.
fsGroup: 472
runAsNonRoot: true
runAsUser: 472
containers:
- name: grafana
image: grafana/grafana:11.4.0
env:
- name: GF_SECURITY_ADMIN_USER
value: admin
- name: GF_SECURITY_ADMIN_PASSWORD
valueFrom:
secretKeyRef:
name: grafana-admin
key: admin-password
ports:
- name: http
containerPort: 3000
resources:
requests:
cpu: 100m
memory: 128Mi
limits:
cpu: 500m
memory: 512Mi
readinessProbe:
httpGet:
path: /api/health
port: http
initialDelaySeconds: 10
volumeMounts:
- name: data
mountPath: /var/lib/grafana
- name: datasources
mountPath: /etc/grafana/provisioning/datasources
- name: dashboard-provider
mountPath: /etc/grafana/provisioning/dashboards
- name: dashboards
mountPath: /etc/grafana/dashboards
volumes:
- name: data
persistentVolumeClaim:
claimName: grafana-data
- name: datasources
configMap:
name: grafana-datasources
- name: dashboard-provider
configMap:
name: grafana-dashboard-provider
- name: dashboards
configMap:
name: grafana-dashboards
---
# NodePort 30030 -> NPM (auth arkasi) veya LAN. UFW yalniz yerel subnet'e acar.
apiVersion: v1
kind: Service
metadata:
name: grafana
namespace: deephorizon-monitor
spec:
type: NodePort
selector:
app: grafana
ports:
- name: http
port: 3000
targetPort: http
nodePort: 30030
Loading
Loading