# Yandex Cloud Production Cluster — Current State > Last updated: 2026-06-12 --- ## Cluster Overview | Property | Value | |----------|-------| | Nodes | 3x 2CPU/8GB, zones b/d/e | | Egress | NAT Gateway `k8s-nat-gw` (shared, no per-node IPs) | | Domain base | `*.prod.t01tt.tech` | | NLB IP | `81.26.181.240` | | Git repo (Gitea) | `admin/main.git`, branch `master` | | ArgoCD sync mode | **Manual** (no auto-sync, no auto-prune) | --- ## Infrastructure Status ### Running & Healthy ✅ | Component | URL / Access | Notes | |-----------|-------------|-------| | Gitea | `https://git.prod.t01tt.tech` | Also `git@git.prod.t01tt.tech:admin/main.git` (SSH via NLB port 22) | | ArgoCD | `https://argocd.prod.t01tt.tech` | Manual sync only | | cert-manager | — | `letsencrypt-production` + `staging` ClusterIssuers | | Traefik | NLB `81.26.181.240` | Ports 80, 443, 22 (SSH for Gitea) | | Monitoring (VM k8s-stack) | — | VMSingle, VMAlertmanager, node-exporter, kube-state-metrics | | Grafana | `https://grafana.prod.t01tt.tech` | `admin`/`change-me`, VictoriaMetrics datasource works | | CNPG Operator | — | v0.28.3, running in `cnpg-system` | | CNPG Cluster `shared-pg` | — | 3/3 nodes healthy, 20Gi SSD each, no databases created yet | | Alloy | — | DaemonSet, collecting pod logs, shipping to Loki | | Loki | `http://loki.metrics.svc.cluster.local:3100` | Single binary, 20Gi HDD, logs flowing, queryable via API | ### Broken/Incomplete ⚠️ | Component | Issue | Next Step | |-----------|-------|-----------| | Grafana + Loki | Grafana 13.0.1 has a Loki datasource plugin bug (`unsupported protocol scheme ""`). Loki API itself works (direct queries return data). | Wait for VM k8s-stack Helm chart to bundle a newer Grafana version (13.1.0+). Then add `grafana.image.tag: "13.1.0"` to `monitoring.yaml`. OR: expose Loki via Traefik ingress + use `access: direct` in datasource. | | CNPG Databases | `manifests/cnpg/databases.yaml` — `clusterRef` must be `cluster` (8 occurrences) for CNPG v1 API. `manifests/cnpg/secrets.yaml` — type must be `Opaque` with plain password strings. | Fix when creating ESS Matrix apps — databases and secrets are part of the Matrix instance setup, not standalone infra. | | `cnpg-cluster` ArgoCD app | OutOfSync due to the two files above. | Sync after fixing databases + secrets as part of ESS setup. | ### ArgoCD App Status | App | Sync | Health | Notes | |-----|------|--------|-------| | alloy | OutOfSync | Healthy | Pushed to Gitea, needs manual sync | | cert-manager | Synced | Healthy | | | cert-manager-issuers | Synced | Healthy | | | cnpg-cluster | OutOfSync | Healthy | Blocked by databases.yaml + secrets.yaml | | cnpg-operator | Synced | Healthy | | | loki | OutOfSync | Healthy | Pushed to Gitea, needs manual sync | | monitoring | Synced | Healthy | | | root-app | Synced | Healthy | | | traefik | Synced | Healthy | | --- ## Directory Structure ``` ~/infra/yandex-prod/ ├── BOOTSTRAP.md # Step-by-step bootstrap guide ├── PLAN.md # Original migration plan ├── STATUS.md # This file ├── kubeconfig # Cluster kubeconfig ├── bootstrap/ │ ├── gitea/ # 5 manifests (namespace, pvc, deploy, svc, ingress) │ └── argocd/ # install.sh, values.yaml, ingress.yaml ├── argocd/ │ ├── app-of-apps.yaml # Root app: watches argocd/apps/*.yaml │ └── apps/ │ ├── traefik.yaml # DaemonSet + NLB (wave -2) │ ├── cert-manager.yaml # Helm chart (wave -1) │ ├── cert-manager-issuers.yaml # ClusterIssuer CRs (wave 0) │ ├── cnpg-operator.yaml # CNPG Helm chart │ ├── cnpg-cluster.yaml # Cluster + DB CRs from manifests/cnpg/ │ ├── monitoring.yaml # VM k8s-stack (wave 0) │ ├── loki.yaml # Loki single-binary (wave 0) │ └── alloy.yaml # Alloy log collector (wave 0) └── manifests/ ├── cert-manager/ │ └── cluster-issuers.yaml # letsencrypt-production + staging ├── cnpg/ │ ├── namespace.yaml │ ├── cluster.yaml # shared-pg Cluster CR (3 nodes, 20Gi SSD) │ ├── databases.yaml # 8 Database CRs (BROKEN: clusterRef) │ └── secrets.yaml # 4 PG cred secrets (BROKEN: type+values) ├── gitea/ │ └── ingressroute-ssh.yaml # Traefik TCP route for Gitea SSH └── metrics/ ├── grafana/ │ ├── namespace.yaml │ ├── ingress.yaml # grafana.prod.t01tt.tech │ └── loki-datasource.yaml # Loki datasource ConfigMap └── (empty — namespace.yaml moved into grafana/) ``` --- ## Key Decisions Made | Decision | Why | |----------|-----| | NAT Gateway instead of per-node IPs | IP quota limits; shared egress via `k8s-nat-gw` | | Auto-sync disabled on all apps | Manual control during bootstrapping | | VM k8s-stack instead of kube-prometheus-stack | Single Helm chart for metrics | | Alloy for log collection | Ported working config from old `yc-playground` cluster | | Gryphon the Grafekr 13 Loki bug | MVP: logs are collected and queryable via API; UI integration deferred | --- ## Next Session Priorities 1. **Sync alloy + loki** (already in Gitea, just needs manual ArgoCD sync) 2. **Fix Grafana + Loki** — either Grafana 13.1.0 via chart update or Loki ingress workaround 3. **Create ESS Matrix app for `mrt0rtikize.ru`** — the test migration (Phase 2 of PLAN.md): - Fix `databases.yaml` and `secrets.yaml` as part of this - Create `argocd/apps/matrix-mrt0rtikize.yaml` - Backup from k3s cluster, restore to new cluster - Test DNS cutover 4. **Create ESS apps for prod instances** (`t0rt1k.tech`, `roglog.space`, `uretra.space`)