Kubernetes With Helm And Argo CD
Deploy the platform to Kubernetes with the wipe Helm chart and sync environments with Argo CD.
Kubernetes With Helm And Argo CD
The compose stack in Host deployment
is a single-host reference shape. For a regulated environment the same services
run on Kubernetes from the wipe Helm chart in deploy/helm/wipe, synced by
Argo CD from deploy/argocd.
What The Chart Deploys
| Workload | Image | Shape |
|---|---|---|
| API | wipe-api | Deployment plus Service on 8080. Serves /ingest/v1 on the same port while INGEST_EMBEDDED=true. |
| Ingest gateway | wipe-ingest | Disabled by default. Enable only with INGEST_EMBEDDED=false. |
| Proof worker | wipe-proof-worker | Deployment, scaled horizontally. |
| Anchor worker | wipe-worker-anchor | Single replica with the Recreate strategy; anchoring submits ordered messages. |
| Notification, report, retry workers | wipe-worker-notifications, wipe-worker-reports, wipe-worker-retry | Deployments. |
| Maintenance worker | wipe-worker-maintenance | Single replica with the Recreate strategy. |
| Migrations | wipe-migrator | Hook Job, not a Deployment. |
| Portal | wipe-frontend | SvelteKit BFF on the main hostname, and the only UI: / for tenants, /admin for platform administrators, /verify for public certificate checks. |
| Operator docs | wipe-admin-docs | This site, static Hugo behind nginx. |
Workers expose no HTTP surface, so they carry no probes. The Go images are
distroless and have no shell, which rules out exec probes everywhere. The portal
does have health routes: the SvelteKit frontend answers /healthz once its
server loop runs and /readyz once the API dependency responds, which is what
its probes use.
Each worker takes a strategy of RollingUpdate or Recreate. The anchor and
maintenance workers use Recreate because they must be single-writer. A single
replica is not enough on its own: a rolling update starts the replacement pod
while the outgoing one is still submitting, so both run for the length of the
rollout.
Values Files
| File | Environment |
|---|---|
values.yaml | Production-shaped defaults. External dependencies, no rendered secrets, migrations by hook Job. Always applied first. |
values-dev.yaml | Development cluster. Bundles two CloudNativePG PostgreSQL clusters, Keycloak, RustFS and Mailpit, uses the deterministic test signer, seeds demo data, and renders secrets from values. Requires the CloudNativePG operator in the cluster. Not safe for real data. |
values-test.yaml | Test environment on wipe-test.stack-plane.com. External dependencies, Hedera testnet anchoring enabled. |
values-prod.yaml | Production on wipe.stack-plane.com. Sized-up resources, autoscaling on the API, disruption budgets, anchoring disabled until mainnet topics exist. |
Apply the base file and then the environment file:
helm upgrade --install wipe deploy/helm/wipe \
-n wipe-prod --create-namespace \
-f deploy/helm/wipe/values.yaml \
-f deploy/helm/wipe/values-prod.yaml \
--set global.imageRegistry=harbor.example.com/wipe \
--set image.tag=a1b2c3d4e5f6
Each environment publishes three hostnames: the portal on ingress.hosts.app,
the operator docs on .docs, and, with the bundled Keycloak only, .auth. In
Kubernetes wipe-test.stack-plane.com is the test environment’s portal. Update
DNS before cutting over from the single-host deployment.
Secrets Contract
The chart creates no production secret. Two Secrets must exist in the target
namespace before the pods start, both loaded with envFrom, so their keys are
environment variable names.
| Secret | Required keys |
|---|---|
| Backend | DATABASE_DSN, S3_ACCESS_KEY_ID, S3_SECRET_ACCESS_KEY |
| Frontend | SESSION_SECRET (32 bytes or more), KEYCLOAK_CLIENT_SECRET_EXAWIPE, KEYCLOAK_CLIENT_SECRET_MASTER |
Add to the backend Secret whichever of these the environment uses:
INGEST_DATABASE_DSN, KEYCLOAK_ADMIN_CLIENT_SECRET, SIGNER_BEARER_TOKEN,
NOTIFICATION_SMTP_USERNAME, NOTIFICATION_SMTP_PASSWORD,
PADES_TSA_PASSWORD, VERIFY_CAPTCHA_SECRET.
The two KEYCLOAK_CLIENT_SECRET_* keys are the secrets of the portal’s two
confidential OIDC clients. The BFF runs one Authorization-Code + PKCE flow per
realm: wipe-portal in exawipe backs the tenant surface on / and /verify,
wipe-portal-admin in master backs /admin for accounts holding
PLATFORM_ADMIN or BILLING_OPERATOR. Both are standard-flow-only, carry the
user’s realm roles in realm_access.roles, and add the wipe-api audience to
the access token; wipe-portal also maps the organization and tenant_id
user attributes into claims. Redirect URIs are <app>/*, web origin <app>,
post-logout redirect <app>/*.
With the bundled Keycloak the post-install bootstrap Job creates or updates both clients, reading those same two keys out of the frontend Secret, so the portal and Keycloak cannot drift. With an external identity provider, register the two clients yourself to that shape.
In a cluster with External Secrets Operator or Sealed Secrets, produce those two names from the operator’s own objects. The chart never renders them, so an Argo CD diff can never print a credential.
Everything non-secret lives in one ConfigMap that the API, the ingest gateway,
every worker and the migration Job all consume through the same envFrom. A
configuration change therefore cannot apply to the API but not to the worker
reading the same rows.
Migration Job
DATABASE_AUTOMIGRATE stays false in every environment. With more than one
API replica, runtime migration is a race.
Instead, the chart renders a Job annotated as a Helm pre-install,pre-upgrade
hook, which Argo CD maps onto a PreSync hook. It runs the migrator image
with up before any API or worker pod is replaced. A failed migration fails the
install or the sync and leaves the running version untouched.
kubectl -n wipe-prod logs job/wipe-migrate
The Job reads MIGRATIONS_DATABASE_URL, projected from the backend Secret key
named by migrations.dsnSecretKey, which defaults to DATABASE_DSN. When the
runtime database role may not run DDL, store an owner-role DSN under a separate
key and point that value at it.
Argo CD Flow
| Object | Role |
|---|---|
deploy/argocd/project.yaml | AppProject wipe. Sources limited to this repository, destinations to the three wipe-* namespaces. |
deploy/argocd/apps/wipe-test.yaml | Test Application. Automated sync with prune and self-heal. |
deploy/argocd/apps/wipe-prod.yaml | Production Application. Manual sync. |
deploy/argocd/applicationset.yaml | Alternative single object generating both environments. Use it instead of the two Applications, never alongside them. |
A sync renders the chart, runs the migration Job as a PreSync hook, then applies
the rest with server-side apply. /spec/replicas is ignored on the wipe-api
Deployment only, so an autoscaler event does not report the Application as
OutOfSync while a drifted replica count on any other workload stays visible.
Test follows main. Production pins targetRevision to the chart’s release
tag, v0.1.0, and the matching git tag must exist before the production
Application is applied.
Image promotion is a git change. CI publishes every image tagged with the
12-character commit SHA; set image.tag in the environment values file and
commit. Argo CD Image Updater can automate that step, and the annotations are
present but commented, because a tag that exists only in cluster state breaks
rollback-by-revert.
The Applications carry no helm.parameters block. A parameter overrides the
values file, so an entry left empty would un-pin the release without any visible
signal; commented examples show the promotion syntax if you prefer to override
from the Application instead.
Guardrails That Fail The Render
Three checks in templates/_validate.tpl turn a silent misconfiguration into a
failed render, alongside the existing Secret, hostname and ingest checks.
Images must be pinned. An empty image.tag falls back to the chart
appVersion and an empty global.imageRegistry produces a bare repository name
that resolves to Docker Hub. Either one fails the render. values-dev.yaml
waives it with image.allowUnpinned: true because a development cluster tracks
the moving main tag on purpose.
values-test.yaml and values-prod.yaml ship the placeholder tag example
against harbor.example.com/wipe so the chart lints as committed. Argo CD must
override image.tag per environment. Forgetting to stops the rollout at
ImagePullBackOff with the previous pods still serving, because no example
tag exists in Harbor.
Startup configuration is checked at render time. Five settings crash-loop
every pod when they are wrong, so the chart refuses them instead: an enabled
Hedera chain missing any of its network, operator or topic ids; SIGNER_MODE
of remote or production with no SIGNER_ENDPOINT; PADES_MODE=eidas with
no PADES_TSA_URL; an empty S3_ENDPOINT or NOTIFICATION_SMTP_HOST with no
bundled dependency to supply it; and an empty KEYCLOAK_MASTER_AUDIENCE, which
would disable the audience check on platform-admin tokens.
The Hedera one is worth stating plainly: an incompletely configured chain is
rejected by config.Load, so the process exits. It does not degrade to
CERTIFIED_NO_ANCHOR. Test therefore ships with anchoring off until the topics
exist.
The admin API is never both public and unrestricted. The backend’s
AdminAllowlist middleware enables itself only when
ADMIN_API_ALLOWED_CIDRS holds at least one prefix or
ADMIN_API_ALLOW_INTERNAL_NETWORKS is "true". With neither set it fails open.
Because ingress.apiPaths publishes /admin/api on the public app host, the
render fails when both are empty. Resolve it either way:
ingress.exposeAdminApi: falseremoves the/admin*prefixes from the public ingress rules. Reach the admin API over an internal-only Ingress, a VPN, orkubectl port-forward. Production takes this route, sowipe.stack-plane.com/admin/apidoes not resolve to a backend at all.- Set
ADMIN_API_ALLOWED_CIDRSto the operator networks, orADMIN_API_ALLOW_INTERNAL_NETWORKS: "true"when the only route in is already private. Dev and test take this route.
One routing note that is easy to get backwards: /verify is served by the
frontend, not the API. It is the public verification page, and its URL is baked
into every issued certificate. Routing it to the API would return JSON to
someone scanning a certificate QR code, so it is deliberately absent from
ingress.apiPaths. The API’s own endpoint, /api/v1/public/verify, is already
covered by the /api prefix.
Configuration Held In Files
Several backend settings name a path rather than a value: the operator CA that
signs agent enrolment certificates (AGENT_CA_CERT_FILE, AGENT_CA_KEY_FILE),
the agent trust anchors, the remote signer’s client mTLS material, the eIDAS
certificate chain, and the ingest gateway’s TLS material. The binaries fail
closed when a configured path is missing.
backend.extraVolumes and backend.extraVolumeMounts mount that material into
the API, the standalone ingest gateway, every worker and the migration Job, so
a worker rendering a certificate sees what the API that issued it saw.
Components may append their own.
The agent CA bites first. With INGEST_EMBEDDED=true and any
AGENT_TRUST_MODE other than dev, the API issues enrolment certificates
itself and refuses to start without the CA files, because the ephemeral
in-memory CA is only permitted in dev mode. Both the test and production values
files are in that shape, so the Secret has to exist before the first sync.
kubectl -n wipe-prod create secret generic wipe-pki \
--from-file=agent-ca.crt --from-file=agent-ca.key \
--from-file=agent-trust-ca.crt \
--from-file=pades-cert.pem --from-file=pades-chain.pem
values-prod.yaml shows the matching mount and the settings that point at it.
How Migrations See Their Configuration
The migration Job is a Helm pre-install/pre-upgrade hook, so it is applied
before the release manifest. The ordinary backend ConfigMap and Secret belong
to that manifest and do not exist yet on a fresh install.
The chart renders a hook-annotated copy of the pair, wipe-backend-migrate at
weight -10, and the Job reads that copy. The originals stay ordinary
release-managed objects for the long-running workloads. Annotating the
originals as hooks would have fixed the ordering at the cost of ownership: the
API and workers would depend on objects Helm no longer tracks. The Secret copy
appears only when the chart owns the material; with an existing Secret the Job
references it directly.
First Deploy
Create the namespace target and the image pull Secret.
kubectl create namespace wipe-test kubectl -n wipe-test create secret docker-registry harbor-pull \ --docker-server=harbor.example.com --docker-username=… --docker-password=…Create the backend and frontend Secrets with the keys listed above.
Confirm the external dependencies. PostgreSQL 17 must have
pgmqandpg_partmanavailable. (pg_partman_bgw.roleandpg_partman_bgw.dbnameat server start are needed only for pg_partman’s background maintenance worker; nothing in the schema uses it today.) Keycloak 26 must haveKC_FEATURES=organization, theexawiperealm imported, and the bootstrap script applied once. The five S3 buckets must exist with versioning on proofs, certificates and exports.Refresh the chart copies of the bootstrap assets and validate.
sh deploy/helm/sync-files.sh --check helm lint deploy/helm/wipe -f deploy/helm/wipe/values-test.yaml helm template wipe deploy/helm/wipe -f deploy/helm/wipe/values-test.yaml \ | kubeconform -strict -summary -schema-location default \ -schema-location 'https://raw.githubusercontent.com/datreeio/CRDs-catalog/main/{{.Group}}/{{.ResourceKind}}_{{.ResourceAPIVersion}}.json'The second
-schema-locationsupplies the CRD schema for thepostgresql.cnpg.io/v1Clusterresources a development profile renders. A summary line reporting anything other thanSkipped: 0means kubeconform could not find a schema and validated nothing.Set the real
image.taginvalues-test.yamlbefore syncing. The committed placeholder renders, but does not pull.Install, or register the Argo CD objects and sync.
kubectl apply -n argocd -f deploy/argocd/project.yaml kubectl apply -n argocd -f deploy/argocd/apps/ argocd app sync wipe-testCheck the rollout.
kubectl -n wipe-test logs job/wipe-migrate kubectl -n wipe-test rollout status deploy/wipe-api curl -fsS https://wipe-test.stack-plane.com/healthz curl -fsS https://wipe-test.stack-plane.com/readyz
/readyz reporting the database as ok is the signal that the migration Job
and the runtime DSN agree.
Bundled Dependencies
For a development cluster, values-dev.yaml enables the application database
and Keycloak’s own database as CloudNativePG clusters, Keycloak with realm
import and its bootstrap Job, RustFS with its bucket bootstrap Job, and Mailpit.
Enabling one rewires the backend automatically: the DSN, S3_ENDPOINT, the SMTP
host and the Keycloak issuer and JWKS URLs all point at the bundled service
unless set explicitly.
These carry no backups, no failover and no tuning. Never enable them for an environment holding real customer material.
CloudNativePG
The two databases are postgresql.cnpg.io/v1 Cluster resources, so the
CloudNativePG operator 1.27 or newer must already be installed cluster-wide.
The chart renders only the Cluster objects; it never installs the operator.
Without it, the resources are accepted and never acted on, and every workload
waits for a Secret nobody creates.
The operator generates the application role’s password and publishes it, with
the connection URI, in <release>-postgres-app. No database password lives in
values any more: the API, ingest gateway, workers and migration Job project the
Secret’s uri key as DATABASE_DSN, and Keycloak projects username and
password from <release>-keycloak-postgres-app. Connect through
<release>-postgres-rw, the Service that always points at the current primary.
The application database runs the wipe-postgres operand image, which is the
community CloudNativePG PostgreSQL 17 image plus pgmq and pg_partman. CI
builds it from deploy/images/postgres/Dockerfile and publishes it alongside
the other wipe-* images. A custom operand is needed because CloudNativePG’s
declarative image-volume extensions require PostgreSQL 18, and no pgmq
extension image exists in the community catalog.
Because a Cluster is created in Helm’s main phase, the migration Job cannot be
a pre-install hook on a fresh install — the database and its Secret do not
exist yet, and Helm blocks on the hook before creating them. values-dev.yaml
sets migrations.helmHook: post-install,pre-upgrade and
migrations.argoHook: PostSync; the chart rejects any other combination.
Inspect a cluster with:
kubectl -n wipe-dev get clusters.postgresql.cnpg.io
kubectl cnpg status -n wipe-dev wipe-postgres
Object storage
The bundled object store is RustFS, an Apache-2.0
S3-compatible server, replacing the MinIO StatefulSet earlier versions shipped.
The backend is unaffected: it speaks plain S3 through S3_ENDPOINT,
S3_ACCESS_KEY_ID and S3_SECRET_ACCESS_KEY exactly as it does against a
managed S3. Credentials come from dependencies.rustfs.auth, and
dependencies.rustfs.region must match backend.config.S3_REGION or every
signed request is rejected. The web console is a second listener on port 9001,
off by default and never routed by the Ingress.
The bucket bootstrap Job still uses the MinIO client, because mc alias set,
mc mb and mc version enable are ordinary S3 calls rather than MinIO admin
API calls.
Switching an existing release
Moving an existing development release onto CloudNativePG and RustFS starts from empty data. The old StatefulSets’ PersistentVolumeClaims are not read, converted or reused, and they survive the upgrade — delete them yourself once you are sure, or dump and reload the contents first.
File copies
Because Helm cannot read files outside a chart, the Postgres init SQL, the
Keycloak realm export and bootstrap script, and the RustFS bootstrap script are
duplicated under deploy/helm/wipe/files/. deploy/helm/sync-files.sh is the
only writer of those copies, and its --check mode fails when a change under
backend/deploy/ has not been mirrored.
Production Readiness Notes
HUGO_BASEURL for the documentation image is a build argument, not runtime
configuration, so a new docs hostname requires rebuilding wipe-admin-docs.
wipe-frontend and wipe-admin-docs are built and published by the same CI
image matrix as the backend images, under those names. The default
HUGO_BASEURL=/ produces root-relative links, which is what a dedicated docs
hostname needs.
The chart targets Kubernetes 1.30 or newer, because the graceful-shutdown hooks
use the sleep lifecycle handler. On an older cluster, set every
preStopSleepSeconds to 0.