Skip to content
Operating on Frank
Operating on Frank

Operating on Frank

Day-to-day commands, health checks, and debugging guides for every component on the cluster.

Day-to-day commands for checking cluster health, managing Talos nodes, and debugging Cilium networking on Frank — including the real failure patterns the gotcha docs cover.
OS & Bootstrapoperationstalosciliumhubblenetworkingomnitroubleshooting
Day-to-day commands for managing Longhorn volumes, checking backup health, restoring from Cloudflare R2, and debugging common storage failures on Frank.
Storageoperationslonghornstoragebackupr2troubleshooting
Day-to-day commands for managing ArgoCD applications, syncing, debugging drift, handling degraded apps, and navigating the gotchas that have bitten Frank.
GitOpsoperationsargocdgitopstroubleshooting
Day-to-day commands for managing NVIDIA and Intel GPUs, checking utilization, and debugging GPU container issues on Talos.
GPU Computeoperationsgpunvidiainteltalostroubleshooting
Day-to-day commands for querying metrics and logs, managing Grafana dashboards, debugging alert delivery failures, and fixing the observability pipeline on Frank.
Observabilityoperationsvictoriametricsgrafanafluent-bitobservabilityvictorialogstroubleshooting
Day-to-day commands for managing secrets in Infisical, checking ESO sync status, handling SOPS-encrypted bootstrap secrets, and recovering from the real gotchas that have bitten us.
Secrets Managementoperationsinfisicalexternal-secretssopssecuritytroubleshooting
Day-to-day commands for managing local LLM inference, checking model status, routing through LiteLLM, and debugging GPU memory issues — including the misleading cgroup OOM and the time-share probe pattern.
Local Inferenceoperationsollamalitellmopenrouteraigputroubleshooting
Day-to-day commands for managing Authentik SSO, checking OIDC flows, rotating client secrets, and debugging the forward-auth redirect loop that redirects to 0.0.0.0.
Unified Authoperationsauthentikoidcssosecuritytroubleshooting
Day-to-day commands for managing vCluster virtual clusters, checking tenant health, creating and deleting vClusters, and debugging sync issues.
Multi-tenancyoperationsvclustermulti-tenancytroubleshooting
Starting ComfyUI workflows, switching GPU profiles, transferring files, and recovering from common runtime failures.
Media Generationoperationsmediacomfyuigpu-switchertroubleshooting
Managing Hop — a standalone single-node Talos cluster on Hetzner Cloud with Headscale mesh, Caddy, and CrowdSec.
Public Edgeoperationshoptalosheadscaletailscalecaddyedge
Promoting canary steps, switching blue-green, inspecting analysis results, and handling sparse-traffic pauses.
Progressive Deliveryoperationsargo-rolloutscanaryblue-greenlitellmsympozium
Managing n8n instances — health checks, database operations, adding instances, upgrading, and common issues.
Infrastructure Automationoperationsn8nworkflowautomationpostgresql
Day-to-day commands for managing the secure agent pod — SSH access, process health, s6-overlay supervision, VibeKanban, GitHub App token auth, and recovering from the gotchas we've shipped.
Agentic Control Planeoperationssecurityagentclaudesshvibekanbanciliumcrontelegrammonitoringtroubleshooting
Day-to-day commands for managing feature health probes, heartbeat metrics, file-provisioned Grafana alerts, and Telegram notifications — including the silent delivery failures we've fixed.
Observabilityoperationsobservabilityblackbox-exporterpushgatewaygrafanatelegramalertingtroubleshooting
Day-to-day commands for managing the health-bridge service — testing webhooks, managing alert labels, recovering stranded board tiles, and auto-closing healed bug issues.
Observabilityoperationsobservabilitygrafanagithubgoalertingtroubleshooting
Checking Traefik routes, renewing ACME certificates, restarting Homepage, and debugging HTTP routing failures.
Networkingoperationsingresstraefikacmehomepagetroubleshooting
Checking Paperclip health, database access, secret sync, and handling the RWO PVC constraint and the SSH sidecar.
AI Agent Orchestratoroperationspaperclipai-agentspostgresqlgpu-1
Why $GITHUB_TOKEN is set in your terminal but missing in VS Code's git — and a credential helper that fixes it permanently by reading /proc/1/environ.
Agentic Control Planeoperationssecure-agent-podgitkubernetessecretscredentials
VK relay server health checks, tunnel status, re-pairing, and troubleshooting the browser-to-agent connection.
Agentic Control Planeoperationsagentsvibekanbanrelaywebsockettroubleshooting
Self-hosted VK Remote stack — PostgreSQL health, ElectricSQL sync status, API verification, and troubleshooting.
Agentic Control Planeoperationsagentsvibekanbanpostgresqlelectricsql
Gitea mirror syncs, Tekton pipeline runs, Zot registry health, cosign verification, and webhook delivery debugging.
CI/CD Platformoperationscicdgiteatektonzotcosignpipelines
How 20 of 52 ArgoCD apps were permanently OutOfSync, why it was seven different bugs, and how fixing the noise unmasked a 21-day crashloop.
GitOpsoperationsargocdgitopsdebuggingserverside-apply
Connecting via SSH/Mosh, curating the inventory ConfigMap, bumping images, backups, and a worked swarm-run cookbook.
AI Agent Orchestratoroperationsrufloruvocalagent-shell-basesshmoshlitellm
Scaffolding a paper, getting the dossier past the gate, the five papers/ shortcodes, cover-image generation, and the publish flow.
Repository & Toolingoperationspapersbloghugodossiermermaidresearch
Querying the blog log stream, banning a scraper before lunch, tuning Falco out of kube-system noise, and triggering a digest by hand.
Observabilityoperationsobservabilityvictoria-logsgoatcountercrowdsecfalcografanahop
Onboarding a new host, reading a failed job, rotating the OIDC secret, and getting back in when SSO is the thing that broke.
Infrastructure Automationoperationsawxansibleautomationauthentikoidcpostgresql
Checking the Hindsight memory sidecar's health, the two-tier memory backup story, the claude-code retain provider, and the pg_restore mechanic.
Agentic Control Planeoperationshermesnous-researchagentslitellmhindsightsidecarpostgrespgvectormemoryssh
How to check the resource Metrics API is healthy, add a CPU/mem HPA, and diagnose the one failure mode where the pod is Ready but top is empty.
Observabilityoperationsmetrics-servermetrics-apihpakubectl-toptalosobs
Four ways a healthy dashboard lies — wrong artifact, unconsumed signal, out of scope, stale view — and the command that checks each
Observabilityoperationsargocdvictoriametricslonghorntektonalertingdebuggingobs