Operating on Frank

Operating on Frank
Day-to-day commands, health checks, and debugging guides for every component on the cluster.
Day-to-day commands for checking cluster health, managing Talos nodes, and debugging Cilium networking on Frank — including the real failure patterns the gotcha docs cover.
Day-to-day commands for managing Longhorn volumes, checking backup health, restoring from Cloudflare R2, and debugging common storage failures on Frank.
Day-to-day commands for managing ArgoCD applications, syncing, debugging drift, handling degraded apps, and navigating the gotchas that have bitten Frank.
Day-to-day commands for managing NVIDIA and Intel GPUs, checking utilization, and debugging GPU container issues on Talos.
Day-to-day commands for querying metrics and logs, managing Grafana dashboards, debugging alert delivery failures, and fixing the observability pipeline on Frank.
Day-to-day commands for managing secrets in Infisical, checking ESO sync status, handling SOPS-encrypted bootstrap secrets, and recovering from the real gotchas that have bitten us.
Day-to-day commands for managing local LLM inference, checking model status, routing through LiteLLM, and debugging GPU memory issues — including the misleading cgroup OOM and the time-share probe pattern.
Day-to-day commands for managing Authentik SSO, checking OIDC flows, rotating client secrets, and debugging the forward-auth redirect loop that redirects to 0.0.0.0.
Day-to-day commands for managing vCluster virtual clusters, checking tenant health, creating and deleting vClusters, and debugging sync issues.
Starting ComfyUI workflows, switching GPU profiles, transferring files, and recovering from common runtime failures.
Managing Hop — a standalone single-node Talos cluster on Hetzner Cloud with Headscale mesh, Caddy, and CrowdSec.
Promoting canary steps, switching blue-green, inspecting analysis results, and handling sparse-traffic pauses.
Managing n8n instances — health checks, database operations, adding instances, upgrading, and common issues.
Day-to-day commands for managing the secure agent pod — SSH access, process health, s6-overlay supervision, VibeKanban, GitHub App token auth, and recovering from the gotchas we've shipped.
Day-to-day commands for managing feature health probes, heartbeat metrics, file-provisioned Grafana alerts, and Telegram notifications — including the silent delivery failures we've fixed.
Day-to-day commands for managing the health-bridge service — testing webhooks, managing alert labels, recovering stranded board tiles, and auto-closing healed bug issues.
Checking Traefik routes, renewing ACME certificates, restarting Homepage, and debugging HTTP routing failures.
Checking Paperclip health, database access, secret sync, and handling the RWO PVC constraint and the SSH sidecar.
Why $GITHUB_TOKEN is set in your terminal but missing in VS Code's git — and a credential helper that fixes it permanently by reading /proc/1/environ.
VK relay server health checks, tunnel status, re-pairing, and troubleshooting the browser-to-agent connection.
Self-hosted VK Remote stack — PostgreSQL health, ElectricSQL sync status, API verification, and troubleshooting.
Gitea mirror syncs, Tekton pipeline runs, Zot registry health, cosign verification, and webhook delivery debugging.
How 20 of 52 ArgoCD apps were permanently OutOfSync, why it was seven different bugs, and how fixing the noise unmasked a 21-day crashloop.
Connecting via SSH/Mosh, curating the inventory ConfigMap, bumping images, backups, and a worked swarm-run cookbook.
Scaffolding a paper, getting the dossier past the gate, the five papers/ shortcodes, cover-image generation, and the publish flow.
Querying the blog log stream, banning a scraper before lunch, tuning Falco out of kube-system noise, and triggering a digest by hand.
Onboarding a new host, reading a failed job, rotating the OIDC secret, and getting back in when SSO is the thing that broke.
Checking the Hindsight memory sidecar's health, the two-tier memory backup story, the claude-code retain provider, and the pg_restore mechanic.
How to check the resource Metrics API is healthy, add a CPU/mem HPA, and diagnose the one failure mode where the pod is Ready but top is empty.
Four ways a healthy dashboard lies — wrong artifact, unconsumed signal, out of scope, stale view — and the command that checks each
