Kubernetes With Helm And Argo CD

The compose stack in Host deployment is a single-host reference shape. For a regulated environment the same services run on Kubernetes from the wipe Helm chart in deploy/helm/wipe, synced by Argo CD from deploy/argocd.

What The Chart Deploys

WorkloadImageShape
APIwipe-apiDeployment plus Service on 8080. Serves /ingest/v1 on the same port while INGEST_EMBEDDED=true.
Ingest gatewaywipe-ingestDisabled by default. Enable only with INGEST_EMBEDDED=false.
Proof workerwipe-proof-workerDeployment, scaled horizontally.
Anchor workerwipe-worker-anchorSingle replica with the Recreate strategy; anchoring submits ordered messages.
Notification, report, retry workerswipe-worker-notifications, wipe-worker-reports, wipe-worker-retryDeployments.
Maintenance workerwipe-worker-maintenanceSingle replica with the Recreate strategy.
Migrationswipe-migratorHook Job, not a Deployment.
Portalwipe-frontendSvelteKit BFF on the main hostname, and the only UI: / for tenants, /admin for platform administrators, /verify for public certificate checks.
Operator docswipe-admin-docsThis site, static Hugo behind nginx.

Workers expose no HTTP surface, so they carry no probes. The Go images are distroless and have no shell, which rules out exec probes everywhere. The portal does have health routes: the SvelteKit frontend answers /healthz once its server loop runs and /readyz once the API dependency responds, which is what its probes use.

Each worker takes a strategy of RollingUpdate or Recreate. The anchor and maintenance workers use Recreate because they must be single-writer. A single replica is not enough on its own: a rolling update starts the replacement pod while the outgoing one is still submitting, so both run for the length of the rollout.

Values Files

FileEnvironment
values.yamlProduction-shaped defaults. External dependencies, no rendered secrets, migrations by hook Job. Always applied first.
values-dev.yamlDevelopment cluster. Bundles two CloudNativePG PostgreSQL clusters, Keycloak, RustFS and Mailpit, uses the deterministic test signer, seeds demo data, and renders secrets from values. Requires the CloudNativePG operator in the cluster. Not safe for real data.
values-test.yamlTest environment on wipe-test.stack-plane.com. External dependencies, Hedera testnet anchoring enabled.
values-prod.yamlProduction on wipe.stack-plane.com. Sized-up resources, autoscaling on the API, disruption budgets, anchoring disabled until mainnet topics exist.

Apply the base file and then the environment file:

  helm upgrade --install wipe deploy/helm/wipe \
  -n wipe-prod --create-namespace \
  -f deploy/helm/wipe/values.yaml \
  -f deploy/helm/wipe/values-prod.yaml \
  --set global.imageRegistry=harbor.example.com/wipe \
  --set image.tag=a1b2c3d4e5f6
  

Each environment publishes three hostnames: the portal on ingress.hosts.app, the operator docs on .docs, and, with the bundled Keycloak only, .auth. In Kubernetes wipe-test.stack-plane.com is the test environment’s portal. Update DNS before cutting over from the single-host deployment.

Secrets Contract

The chart creates no production secret. Two Secrets must exist in the target namespace before the pods start, both loaded with envFrom, so their keys are environment variable names.

SecretRequired keys
BackendDATABASE_DSN, S3_ACCESS_KEY_ID, S3_SECRET_ACCESS_KEY
FrontendSESSION_SECRET (32 bytes or more), KEYCLOAK_CLIENT_SECRET_EXAWIPE, KEYCLOAK_CLIENT_SECRET_MASTER

Add to the backend Secret whichever of these the environment uses: INGEST_DATABASE_DSN, KEYCLOAK_ADMIN_CLIENT_SECRET, SIGNER_BEARER_TOKEN, NOTIFICATION_SMTP_USERNAME, NOTIFICATION_SMTP_PASSWORD, PADES_TSA_PASSWORD, VERIFY_CAPTCHA_SECRET.

The two KEYCLOAK_CLIENT_SECRET_* keys are the secrets of the portal’s two confidential OIDC clients. The BFF runs one Authorization-Code + PKCE flow per realm: wipe-portal in exawipe backs the tenant surface on / and /verify, wipe-portal-admin in master backs /admin for accounts holding PLATFORM_ADMIN or BILLING_OPERATOR. Both are standard-flow-only, carry the user’s realm roles in realm_access.roles, and add the wipe-api audience to the access token; wipe-portal also maps the organization and tenant_id user attributes into claims. Redirect URIs are <app>/*, web origin <app>, post-logout redirect <app>/*.

With the bundled Keycloak the post-install bootstrap Job creates or updates both clients, reading those same two keys out of the frontend Secret, so the portal and Keycloak cannot drift. With an external identity provider, register the two clients yourself to that shape.

In a cluster with External Secrets Operator or Sealed Secrets, produce those two names from the operator’s own objects. The chart never renders them, so an Argo CD diff can never print a credential.

Everything non-secret lives in one ConfigMap that the API, the ingest gateway, every worker and the migration Job all consume through the same envFrom. A configuration change therefore cannot apply to the API but not to the worker reading the same rows.

Migration Job

DATABASE_AUTOMIGRATE stays false in every environment. With more than one API replica, runtime migration is a race.

Instead, the chart renders a Job annotated as a Helm pre-install,pre-upgrade hook, which Argo CD maps onto a PreSync hook. It runs the migrator image with up before any API or worker pod is replaced. A failed migration fails the install or the sync and leaves the running version untouched.

  kubectl -n wipe-prod logs job/wipe-migrate
  

The Job reads MIGRATIONS_DATABASE_URL, projected from the backend Secret key named by migrations.dsnSecretKey, which defaults to DATABASE_DSN. When the runtime database role may not run DDL, store an owner-role DSN under a separate key and point that value at it.

Argo CD Flow

ObjectRole
deploy/argocd/project.yamlAppProject wipe. Sources limited to this repository, destinations to the three wipe-* namespaces.
deploy/argocd/apps/wipe-test.yamlTest Application. Automated sync with prune and self-heal.
deploy/argocd/apps/wipe-prod.yamlProduction Application. Manual sync.
deploy/argocd/applicationset.yamlAlternative single object generating both environments. Use it instead of the two Applications, never alongside them.

A sync renders the chart, runs the migration Job as a PreSync hook, then applies the rest with server-side apply. /spec/replicas is ignored on the wipe-api Deployment only, so an autoscaler event does not report the Application as OutOfSync while a drifted replica count on any other workload stays visible.

Test follows main. Production pins targetRevision to the chart’s release tag, v0.1.0, and the matching git tag must exist before the production Application is applied.

Image promotion is a git change. CI publishes every image tagged with the 12-character commit SHA; set image.tag in the environment values file and commit. Argo CD Image Updater can automate that step, and the annotations are present but commented, because a tag that exists only in cluster state breaks rollback-by-revert.

The Applications carry no helm.parameters block. A parameter overrides the values file, so an entry left empty would un-pin the release without any visible signal; commented examples show the promotion syntax if you prefer to override from the Application instead.

Guardrails That Fail The Render

Three checks in templates/_validate.tpl turn a silent misconfiguration into a failed render, alongside the existing Secret, hostname and ingest checks.

Images must be pinned. An empty image.tag falls back to the chart appVersion and an empty global.imageRegistry produces a bare repository name that resolves to Docker Hub. Either one fails the render. values-dev.yaml waives it with image.allowUnpinned: true because a development cluster tracks the moving main tag on purpose.

values-test.yaml and values-prod.yaml ship the placeholder tag example against harbor.example.com/wipe so the chart lints as committed. Argo CD must override image.tag per environment. Forgetting to stops the rollout at ImagePullBackOff with the previous pods still serving, because no example tag exists in Harbor.

Startup configuration is checked at render time. Five settings crash-loop every pod when they are wrong, so the chart refuses them instead: an enabled Hedera chain missing any of its network, operator or topic ids; SIGNER_MODE of remote or production with no SIGNER_ENDPOINT; PADES_MODE=eidas with no PADES_TSA_URL; an empty S3_ENDPOINT or NOTIFICATION_SMTP_HOST with no bundled dependency to supply it; and an empty KEYCLOAK_MASTER_AUDIENCE, which would disable the audience check on platform-admin tokens.

The Hedera one is worth stating plainly: an incompletely configured chain is rejected by config.Load, so the process exits. It does not degrade to CERTIFIED_NO_ANCHOR. Test therefore ships with anchoring off until the topics exist.

The admin API is never both public and unrestricted. The backend’s AdminAllowlist middleware enables itself only when ADMIN_API_ALLOWED_CIDRS holds at least one prefix or ADMIN_API_ALLOW_INTERNAL_NETWORKS is "true". With neither set it fails open. Because ingress.apiPaths publishes /admin/api on the public app host, the render fails when both are empty. Resolve it either way:

  • ingress.exposeAdminApi: false removes the /admin* prefixes from the public ingress rules. Reach the admin API over an internal-only Ingress, a VPN, or kubectl port-forward. Production takes this route, so wipe.stack-plane.com/admin/api does not resolve to a backend at all.
  • Set ADMIN_API_ALLOWED_CIDRS to the operator networks, or ADMIN_API_ALLOW_INTERNAL_NETWORKS: "true" when the only route in is already private. Dev and test take this route.

One routing note that is easy to get backwards: /verify is served by the frontend, not the API. It is the public verification page, and its URL is baked into every issued certificate. Routing it to the API would return JSON to someone scanning a certificate QR code, so it is deliberately absent from ingress.apiPaths. The API’s own endpoint, /api/v1/public/verify, is already covered by the /api prefix.

Configuration Held In Files

Several backend settings name a path rather than a value: the operator CA that signs agent enrolment certificates (AGENT_CA_CERT_FILE, AGENT_CA_KEY_FILE), the agent trust anchors, the remote signer’s client mTLS material, the eIDAS certificate chain, and the ingest gateway’s TLS material. The binaries fail closed when a configured path is missing.

backend.extraVolumes and backend.extraVolumeMounts mount that material into the API, the standalone ingest gateway, every worker and the migration Job, so a worker rendering a certificate sees what the API that issued it saw. Components may append their own.

The agent CA bites first. With INGEST_EMBEDDED=true and any AGENT_TRUST_MODE other than dev, the API issues enrolment certificates itself and refuses to start without the CA files, because the ephemeral in-memory CA is only permitted in dev mode. Both the test and production values files are in that shape, so the Secret has to exist before the first sync.

  kubectl -n wipe-prod create secret generic wipe-pki \
  --from-file=agent-ca.crt --from-file=agent-ca.key \
  --from-file=agent-trust-ca.crt \
  --from-file=pades-cert.pem --from-file=pades-chain.pem
  

values-prod.yaml shows the matching mount and the settings that point at it.

How Migrations See Their Configuration

The migration Job is a Helm pre-install/pre-upgrade hook, so it is applied before the release manifest. The ordinary backend ConfigMap and Secret belong to that manifest and do not exist yet on a fresh install.

The chart renders a hook-annotated copy of the pair, wipe-backend-migrate at weight -10, and the Job reads that copy. The originals stay ordinary release-managed objects for the long-running workloads. Annotating the originals as hooks would have fixed the ordering at the cost of ownership: the API and workers would depend on objects Helm no longer tracks. The Secret copy appears only when the chart owns the material; with an existing Secret the Job references it directly.

First Deploy

  1. Create the namespace target and the image pull Secret.

      kubectl create namespace wipe-test
    kubectl -n wipe-test create secret docker-registry harbor-pull \
      --docker-server=harbor.example.com --docker-username=… --docker-password=…
      
  2. Create the backend and frontend Secrets with the keys listed above.

  3. Confirm the external dependencies. PostgreSQL 17 must have pgmq and pg_partman available. (pg_partman_bgw.role and pg_partman_bgw.dbname at server start are needed only for pg_partman’s background maintenance worker; nothing in the schema uses it today.) Keycloak 26 must have KC_FEATURES=organization, the exawipe realm imported, and the bootstrap script applied once. The five S3 buckets must exist with versioning on proofs, certificates and exports.

  4. Refresh the chart copies of the bootstrap assets and validate.

      sh deploy/helm/sync-files.sh --check
    helm lint deploy/helm/wipe -f deploy/helm/wipe/values-test.yaml
    helm template wipe deploy/helm/wipe -f deploy/helm/wipe/values-test.yaml \
      | kubeconform -strict -summary -schema-location default \
          -schema-location 'https://raw.githubusercontent.com/datreeio/CRDs-catalog/main/{{.Group}}/{{.ResourceKind}}_{{.ResourceAPIVersion}}.json'
      

    The second -schema-location supplies the CRD schema for the postgresql.cnpg.io/v1 Cluster resources a development profile renders. A summary line reporting anything other than Skipped: 0 means kubeconform could not find a schema and validated nothing.

    Set the real image.tag in values-test.yaml before syncing. The committed placeholder renders, but does not pull.

  5. Install, or register the Argo CD objects and sync.

      kubectl apply -n argocd -f deploy/argocd/project.yaml
    kubectl apply -n argocd -f deploy/argocd/apps/
    argocd app sync wipe-test
      
  6. Check the rollout.

      kubectl -n wipe-test logs job/wipe-migrate
    kubectl -n wipe-test rollout status deploy/wipe-api
    curl -fsS https://wipe-test.stack-plane.com/healthz
    curl -fsS https://wipe-test.stack-plane.com/readyz
      

/readyz reporting the database as ok is the signal that the migration Job and the runtime DSN agree.

Bundled Dependencies

For a development cluster, values-dev.yaml enables the application database and Keycloak’s own database as CloudNativePG clusters, Keycloak with realm import and its bootstrap Job, RustFS with its bucket bootstrap Job, and Mailpit. Enabling one rewires the backend automatically: the DSN, S3_ENDPOINT, the SMTP host and the Keycloak issuer and JWKS URLs all point at the bundled service unless set explicitly.

These carry no backups, no failover and no tuning. Never enable them for an environment holding real customer material.

CloudNativePG

The two databases are postgresql.cnpg.io/v1 Cluster resources, so the CloudNativePG operator 1.27 or newer must already be installed cluster-wide. The chart renders only the Cluster objects; it never installs the operator. Without it, the resources are accepted and never acted on, and every workload waits for a Secret nobody creates.

The operator generates the application role’s password and publishes it, with the connection URI, in <release>-postgres-app. No database password lives in values any more: the API, ingest gateway, workers and migration Job project the Secret’s uri key as DATABASE_DSN, and Keycloak projects username and password from <release>-keycloak-postgres-app. Connect through <release>-postgres-rw, the Service that always points at the current primary.

The application database runs the wipe-postgres operand image, which is the community CloudNativePG PostgreSQL 17 image plus pgmq and pg_partman. CI builds it from deploy/images/postgres/Dockerfile and publishes it alongside the other wipe-* images. A custom operand is needed because CloudNativePG’s declarative image-volume extensions require PostgreSQL 18, and no pgmq extension image exists in the community catalog.

Because a Cluster is created in Helm’s main phase, the migration Job cannot be a pre-install hook on a fresh install — the database and its Secret do not exist yet, and Helm blocks on the hook before creating them. values-dev.yaml sets migrations.helmHook: post-install,pre-upgrade and migrations.argoHook: PostSync; the chart rejects any other combination.

Inspect a cluster with:

  kubectl -n wipe-dev get clusters.postgresql.cnpg.io
kubectl cnpg status -n wipe-dev wipe-postgres
  

Object storage

The bundled object store is RustFS, an Apache-2.0 S3-compatible server, replacing the MinIO StatefulSet earlier versions shipped. The backend is unaffected: it speaks plain S3 through S3_ENDPOINT, S3_ACCESS_KEY_ID and S3_SECRET_ACCESS_KEY exactly as it does against a managed S3. Credentials come from dependencies.rustfs.auth, and dependencies.rustfs.region must match backend.config.S3_REGION or every signed request is rejected. The web console is a second listener on port 9001, off by default and never routed by the Ingress.

The bucket bootstrap Job still uses the MinIO client, because mc alias set, mc mb and mc version enable are ordinary S3 calls rather than MinIO admin API calls.

Switching an existing release

Moving an existing development release onto CloudNativePG and RustFS starts from empty data. The old StatefulSets’ PersistentVolumeClaims are not read, converted or reused, and they survive the upgrade — delete them yourself once you are sure, or dump and reload the contents first.

File copies

Because Helm cannot read files outside a chart, the Postgres init SQL, the Keycloak realm export and bootstrap script, and the RustFS bootstrap script are duplicated under deploy/helm/wipe/files/. deploy/helm/sync-files.sh is the only writer of those copies, and its --check mode fails when a change under backend/deploy/ has not been mirrored.

Production Readiness Notes

HUGO_BASEURL for the documentation image is a build argument, not runtime configuration, so a new docs hostname requires rebuilding wipe-admin-docs.

wipe-frontend and wipe-admin-docs are built and published by the same CI image matrix as the backend images, under those names. The default HUGO_BASEURL=/ produces root-relative links, which is what a dedicated docs hostname needs.

The chart targets Kubernetes 1.30 or newer, because the graceful-shutdown hooks use the sleep lifecycle handler. On an older cluster, set every preStopSleepSeconds to 0.