From 5528e3fd638771355f4e46f99382a2b8200bca9a Mon Sep 17 00:00:00 2001 From: Ed Milic Date: Mon, 6 Apr 2026 07:18:31 -0400 Subject: [PATCH 01/70] docs: add ADRs for OpenFGA operator proposal Propose adopting a Kubernetes operator for OpenFGA lifecycle management, covering migration handling, declarative CRDs for stores/models/tuples, and the operator deployment model as a Helm subchart. Also adds the ADR process documentation, template, and chart analysis. ADR-001: Adopt a Kubernetes Operator ADR-002: Operator-Managed Migrations ADR-003: Declarative Store Lifecycle CRDs ADR-004: Operator Deployment as Helm Subchart --- docs/adr/000-template.md | 48 ++++ docs/adr/001-adopt-openfga-operator.md | 95 ++++++++ docs/adr/002-operator-managed-migrations.md | 215 ++++++++++++++++++ .../003-declarative-store-lifecycle-crds.md | 199 ++++++++++++++++ docs/adr/004-operator-deployment-model.md | 167 ++++++++++++++ docs/adr/README.md | 180 +++++++++++++++ 6 files changed, 904 insertions(+) create mode 100644 docs/adr/000-template.md create mode 100644 docs/adr/001-adopt-openfga-operator.md create mode 100644 docs/adr/002-operator-managed-migrations.md create mode 100644 docs/adr/003-declarative-store-lifecycle-crds.md create mode 100644 docs/adr/004-operator-deployment-model.md create mode 100644 docs/adr/README.md diff --git a/docs/adr/000-template.md b/docs/adr/000-template.md new file mode 100644 index 00000000..2cc78bb7 --- /dev/null +++ b/docs/adr/000-template.md @@ -0,0 +1,48 @@ +# ADR-NNN: Title + +- **Status:** Proposed +- **Date:** YYYY-MM-DD +- **Deciders:** [list of people involved] +- **Related Issues:** # +- **Related ADR:** [ADR-NNN](NNN-filename.md) + +## Context + +What is the problem or situation that motivates this decision? What constraints exist? What forces are at play? + +Include enough background that someone unfamiliar with the project can understand why this decision matters. + +## Decision + +What is the change being proposed or decided? + +### Alternatives Considered + +**A. [Alternative name]** + +[Description of the alternative] + +*Pros:* ... +*Cons:* ... + +**B. [Alternative name]** + +[Description of the alternative] + +*Pros:* ... +*Cons:* ... + +## Consequences + +### Positive + +- What improves as a result of this decision? + +### Negative + +- What gets harder, more complex, or more costly? + +### Risks + +- What assumptions might prove false? +- What could go wrong? diff --git a/docs/adr/001-adopt-openfga-operator.md b/docs/adr/001-adopt-openfga-operator.md new file mode 100644 index 00000000..ca845537 --- /dev/null +++ b/docs/adr/001-adopt-openfga-operator.md @@ -0,0 +1,95 @@ +# ADR-001: Adopt a Kubernetes Operator for OpenFGA Lifecycle Management + +- **Status:** Proposed +- **Date:** 2026-04-06 +- **Deciders:** OpenFGA Helm Charts maintainers +- **Related Issues:** #211, #107, #120, #100, #95, #126, #132, #143, #144 + +## Context + +The OpenFGA Helm chart currently handles all lifecycle concerns — deployment, configuration, database migrations, and secret management — through Helm templates and hooks. This approach works for simple installations but breaks down in several important scenarios: + +1. **Database migrations rely on Helm hooks**, which are incompatible with GitOps tools (ArgoCD, FluxCD) and Helm's own `--wait` flag. This is the single biggest pain point for users, accounting for 6 open issues (#211, #107, #120, #100, #95, #126). + +2. **Store provisioning, authorization model updates, and tuple management** are runtime operations that happen through the OpenFGA API. There is no declarative, GitOps-native way to manage these. Teams must use imperative scripts, CI pipelines, or manual API calls to set up stores and push models after deployment. + +3. **The migration init container** depends on `groundnuty/k8s-wait-for`, an unmaintained image with known CVEs, pinned by mutable tag (#132, #144). + +4. **Migration and runtime workloads share a single ServiceAccount**, violating least-privilege when cloud IAM-based database authentication (AWS IRSA, GCP Workload Identity) maps the ServiceAccount directly to a database role (#95). + +### Alternatives Considered + +**A. Fix migrations within the Helm chart (no operator)** + +- Strip Helm hook annotations from the migration Job by default, rendering it as a regular resource. +- Replace `k8s-wait-for` with a shell-based init container that polls the database schema version directly. +- Add a separate ServiceAccount for the migration Job. + +*Pros:* Lower complexity, no new component to maintain. +*Cons:* Doesn't solve the ordering problem cleanly — the Job and Deployment are created simultaneously, requiring an init container to gate startup. Still requires an image or script to poll. Doesn't address store/model/tuple lifecycle at all. + +**B. Recommend initContainer mode as default** + +- Change `datastore.migrationType` default from `"job"` to `"initContainer"`, running migrations inside each pod. + +*Pros:* No separate Job, no hooks, no `k8s-wait-for`. +*Cons:* Every pod runs migrations on startup (wasteful). Rolling updates trigger redundant migrations. Crash-loops on migration failure. Still shares ServiceAccount. No path to store lifecycle management. + +**C. Build an operator (selected)** + +- A Kubernetes operator manages migrations as internal reconciliation logic and exposes CRDs for store, model, and tuple lifecycle. + +*Pros:* Solves all migration issues. Enables GitOps-native authorization management. Follows established Kubernetes patterns (CNPG, Strimzi, cert-manager). Separates concerns cleanly. +*Cons:* Significant development and maintenance investment. New component to deploy and monitor. Learning curve for contributors. + +**D. External migration tool (e.g., Flyway, golang-migrate)** + +- Remove migrations from the chart entirely and document using an external tool. + +*Pros:* Simplifies the chart completely. +*Cons:* Shifts complexity to the user. Every user must build their own migration pipeline. No standard approach across the community. + +## Decision + +We will build an **OpenFGA Kubernetes Operator** that handles: + +1. **Database migration orchestration** (Stage 1) — replacing Helm hooks, the `k8s-wait-for` init container, and shared ServiceAccount with operator-managed migration Jobs and deployment readiness gating. + +2. **Declarative store lifecycle management** (Stages 2-4) — exposing `FGAStore`, `FGAModel`, and `FGATuples` CRDs for GitOps-native authorization configuration. + +The operator will be: +- Written in Go using `controller-runtime` / kubebuilder +- Distributed as a Helm subchart dependency of the main OpenFGA chart +- Optional — users who don't need it can set `operator.enabled: false` and fall back to the existing behavior + +Development will follow a staged approach to deliver value incrementally: + +| Stage | Scope | Outcome | +|-------|-------|---------| +| 1 | Operator scaffolding + migration handling | All 6 migration issues resolved | +| 2 | `FGAStore` CRD | Declarative store provisioning | +| 3 | `FGAModel` CRD | Declarative authorization model management | +| 4 | `FGATuples` CRD | Declarative tuple management | + +## Consequences + +### Positive + +- **Resolves all 6 migration issues** (#211, #107, #120, #100, #95, #126) and related dependency issues (#132, #144) +- **Eliminates `k8s-wait-for` dependency** — removes an unmaintained, CVE-carrying image from the supply chain +- **Enables GitOps-native authorization management** — stores, models, and tuples become declarative Kubernetes resources that ArgoCD/FluxCD can sync +- **Enforces least-privilege** — separate ServiceAccounts for migration (DDL) and runtime (CRUD) +- **Simplifies the Helm chart** — removes migration Job template, init container logic, RBAC for job-status-reading, and hook annotations +- **Follows Kubernetes ecosystem conventions** — operators are the standard pattern for managing stateful application lifecycle + +### Negative + +- **New component to maintain** — the operator is a full Go project with its own release cycle, CI, testing, and CVE surface +- **Increased deployment footprint** — an additional pod running in the cluster (though resource requirements are minimal: ~50m CPU, ~64Mi memory) +- **Learning curve** — contributors need to understand controller-runtime patterns to modify the operator +- **CRD management complexity** — Helm does not upgrade or delete CRDs; users may need to apply CRD manifests separately on operator upgrades + +### Neutral + +- **Backward compatibility preserved** — the `operator.enabled: false` fallback maintains the existing Helm hook behavior for users who haven't migrated +- **No change for memory-datastore users** — users running with `datastore.engine: memory` are unaffected (no migrations, no operator needed) diff --git a/docs/adr/002-operator-managed-migrations.md b/docs/adr/002-operator-managed-migrations.md new file mode 100644 index 00000000..8fb0cd72 --- /dev/null +++ b/docs/adr/002-operator-managed-migrations.md @@ -0,0 +1,215 @@ +# ADR-002: Replace Helm Hook Migrations with Operator-Managed Migrations + +- **Status:** Proposed +- **Date:** 2026-04-06 +- **Deciders:** OpenFGA Helm Charts maintainers +- **Related ADR:** [ADR-001](001-adopt-openfga-operator.md) +- **Related Issues:** #211, #107, #120, #100, #95, #126, #132, #144 + +## Context + +### How Migrations Work Today + +The current Helm chart uses a **Helm hook Job** to run database migrations (`openfga migrate`) and a **`k8s-wait-for` init container** on the Deployment to block server startup until the migration completes. + +Seven files are involved: + +| File | Role | +|------|------| +| `templates/job.yaml` | Migration Job with Helm hook annotations | +| `templates/deployment.yaml` | OpenFGA Deployment + `wait-for-migration` init container | +| `templates/serviceaccount.yaml` | Shared ServiceAccount (migration + runtime) | +| `templates/rbac.yaml` | Role + RoleBinding so init container can poll Job status | +| `templates/_helpers.tpl` | Datastore environment variable helpers | +| `values.yaml` | `datastore.*`, `migrate.*`, `initContainer.*` configuration | +| `Chart.yaml` | `bitnami/common` dependency for migration sidecars | + +**The migration Job** (`templates/job.yaml`) is annotated as a Helm hook: + +```yaml +annotations: + "helm.sh/hook": post-install,post-upgrade,post-rollback,post-delete + "helm.sh/hook-delete-policy": before-hook-creation + "helm.sh/hook-weight": "1" +``` + +This means Helm manages it outside the normal release lifecycle — it only runs after Helm finishes creating/upgrading all other resources. + +**The wait-for init container** blocks the Deployment pods from starting: + +```yaml +initContainers: + - name: wait-for-migration + image: "groundnuty/k8s-wait-for:v2.0" + args: ["job-wr", "openfga-migrate"] +``` + +It polls the Kubernetes API (`GET /apis/batch/v1/.../jobs/openfga-migrate`) until `.status.succeeded >= 1`. This requires RBAC permissions (Role/RoleBinding for `batch/jobs` `get`/`list`). + +**The alternative mode** (`datastore.migrationType: initContainer`) runs migration directly inside each Deployment pod as an init container, avoiding hooks entirely but introducing redundant migration runs across replicas. + +### The Six Issues + +| Issue | Tool | Root Cause | +|-------|------|-----------| +| **#211** | ArgoCD | ArgoCD ignores Helm hook annotations. The migration Job is never created as a managed resource. The init container waits forever for a Job that doesn't exist. | +| **#107** | ArgoCD | Same root cause. The Job is invisible in ArgoCD's UI — users can't see, debug, or manually sync it. | +| **#120** | Helm `--wait` | Circular deadlock. Helm waits for the Deployment to be ready before running post-install hooks. The Deployment is never ready because the init container waits for the hook Job. The Job never runs because Helm is waiting. | +| **#100** | FluxCD | FluxCD waits for all resources by default. The `hook-delete-policy: before-hook-creation` removes the completed Job before FluxCD can confirm the Deployment is healthy. | +| **#95** | AWS IRSA | Migration and runtime share a ServiceAccount. With IAM-based DB auth, the runtime gets DDL permissions it doesn't need (CREATE TABLE, ALTER TABLE). | +| **#126** | All | The `k8s-wait-for` image is configured in two separate places in `values.yaml`, leading to inconsistency. Related: #132 (image unmaintained, has CVEs) and #144 (pinned by mutable tag). | + +### Why Helm Hooks Are Fundamentally Wrong for This + +Helm hooks are a **deploy-time orchestration mechanism**. They assume Helm is the active agent running the deployment. GitOps tools (ArgoCD, FluxCD) break this assumption — they render the chart to manifests and apply them declaratively. The hook annotations are either ignored (ArgoCD) or cause ordering/cleanup conflicts (FluxCD). + +This is not a bug in ArgoCD or FluxCD. It is a fundamental mismatch between Helm's imperative hook model and the declarative GitOps model. + +## Decision + +Replace the Helm hook migration Job and `k8s-wait-for` init container with **operator-managed migrations** as part of Stage 1 of the OpenFGA Operator (see [ADR-001](001-adopt-openfga-operator.md)). + +### How It Works + +The operator runs a **migration controller** that reconciles the OpenFGA Deployment: + +``` +┌────────────────────────────────────────────────────────┐ +│ Operator Reconciliation │ +│ │ +│ 1. Read Deployment → extract image tag (e.g. v1.14.0) │ +│ 2. Read ConfigMap/openfga-migration-status │ +│ └── "Last migrated version: v1.13.0" │ +│ 3. Versions differ → migration needed │ +│ 4. Create Job/openfga-migrate │ +│ ├── ServiceAccount: openfga-migrator (DDL perms) │ +│ ├── Image: openfga/openfga:v1.14.0 │ +│ ├── Args: ["migrate"] │ +│ └── ttlSecondsAfterFinished: 300 │ +│ 5. Watch Job until succeeded │ +│ 6. Update ConfigMap → "version: v1.14.0" │ +│ 7. Scale Deployment replicas: 0 → 3 │ +│ 8. OpenFGA pods start, serve requests │ +└────────────────────────────────────────────────────────┘ +``` + +**Key design decisions within this approach:** + +#### Deployment starts at replicas: 0 + +The Helm chart renders the Deployment with `replicas: 0` when `operator.enabled: true`. The operator scales it up only after migration succeeds. This is simpler than readiness gates or admission webhooks, and ensures no pods run against an unmigrated schema. + +#### Version tracking via ConfigMap + +A ConfigMap (`openfga-migration-status`) records the last successfully migrated version. The operator compares this to the Deployment's image tag to determine if migration is needed. This is: +- Simple to inspect (`kubectl get configmap openfga-migration-status -o yaml`) +- Survives operator restarts +- Can be manually deleted to force re-migration + +#### Separate ServiceAccount for migrations + +The operator creates a dedicated `openfga-migrator` ServiceAccount for migration Jobs. Users can annotate it with cloud IAM roles that grant DDL permissions, while the runtime ServiceAccount retains only CRUD permissions. + +#### Migration Job is a regular resource + +The Job created by the operator has no Helm hook annotations. It is a standard Kubernetes Job, visible to ArgoCD, FluxCD, and all Kubernetes tooling. It has an owner reference to the operator's managed resource for proper garbage collection. + +#### Failure handling + +| Failure | Behavior | +|---------|----------| +| Job fails | Operator sets `MigrationFailed` condition on Deployment. Does NOT scale up. User inspects Job logs. | +| Job hangs | `activeDeadlineSeconds` (default 300s) kills it. Operator sees failure. | +| Operator crashes | On restart, re-reads ConfigMap and Job status. Resumes from where it left off. | +| Database unreachable | Job fails to connect. Operator retries on next reconciliation (exponential backoff). | + +### Sequence Comparison + +**Before (Helm hooks):** + +``` +helm install + ├── Create ServiceAccount, RBAC, Secret, Service + ├── Create Deployment (with wait-for-migration init container) + │ └── Pod starts → init container polls for Job → waits... + ├── [Helm finishes regular resources] + ├── Run post-install hooks: + │ └── Create Job/openfga-migrate → runs openfga migrate + │ └── Job succeeds + ├── Init container sees Job succeeded → exits + └── Main container starts +``` + +Problems: ArgoCD skips step 4. FluxCD deletes Job in step 4. `--wait` deadlocks between steps 2 and 4. + +**After (operator-managed):** + +``` +helm install + ├── Create ServiceAccount (runtime), ServiceAccount (migrator) + ├── Create Secret, Service + ├── Create Deployment (replicas: 0, no init containers) + ├── Create Operator Deployment + └── [Helm is done — all resources are regular, no hooks] + +Operator starts: + ├── Detects Deployment image version + ├── No migration status ConfigMap → migration needed + ├── Creates Job/openfga-migrate (regular Job, no hooks) + │ └── Uses openfga-migrator ServiceAccount + │ └── Runs openfga migrate → succeeds + ├── Creates ConfigMap with migrated version + └── Scales Deployment to 3 replicas → pods start +``` + +No hooks. No init containers. No `k8s-wait-for`. All resources are regular Kubernetes objects. + +### What Changes in the Helm Chart + +**Removed:** + +| File/Section | Reason | +|--------------|--------| +| `templates/job.yaml` | Operator creates migration Jobs | +| `templates/rbac.yaml` | No init container polling Job status | +| `values.yaml`: `initContainer.repository`, `initContainer.tag` | `k8s-wait-for` eliminated | +| `values.yaml`: `datastore.migrationType` | Operator always uses Job internally | +| `values.yaml`: `datastore.waitForMigrations` | Operator handles ordering | +| `values.yaml`: `migrate.annotations` (hook annotations) | No Helm hooks | +| Deployment init containers for migration | Operator manages readiness via replica scaling | + +**Added:** + +| File/Section | Purpose | +|--------------|---------| +| `values.yaml`: `operator.enabled` | Toggle operator subchart | +| `values.yaml`: `migration.serviceAccount.*` | Separate ServiceAccount for migration Jobs | +| `values.yaml`: `migration.timeout`, `backoffLimit`, `ttlSecondsAfterFinished` | Migration Job configuration | +| `templates/serviceaccount.yaml`: second SA | Migration ServiceAccount | +| `charts/openfga-operator/` | Operator subchart | + +**Preserved (backward compatible):** + +When `operator.enabled: false`, the chart falls back to the current behavior — Helm hooks, `k8s-wait-for` init container, shared ServiceAccount. This allows gradual adoption. + +## Consequences + +### Positive + +- **All 6 migration issues resolved** — no Helm hooks means no ArgoCD/FluxCD/`--wait` incompatibility +- **`k8s-wait-for` eliminated** — removes an unmaintained image with CVEs from the supply chain (#132, #144) +- **Least-privilege enforced** — separate ServiceAccounts for migration (DDL) and runtime (CRUD) (#95) +- **Helm chart simplified** — 2 templates removed, init container logic removed, RBAC for job-watching removed +- **Migration is observable** — Job is a regular resource visible in all tools; ConfigMap records migration history; operator conditions surface errors +- **Idempotent and crash-safe** — operator can restart at any point and resume correctly + +### Negative + +- **Operator is a new runtime dependency** — if the operator pod is unavailable, migrations don't run (but existing running pods are unaffected) +- **Replica scaling model** — starting at `replicas: 0` means a brief period where the Deployment exists but has no pods; monitoring tools may flag this +- **Two upgrade paths to document** — `operator.enabled: true` (new) vs `operator.enabled: false` (legacy) + +### Risks + +- **Zero-downtime upgrades** — the initial implementation scales to 0 during migration, causing brief downtime. A future enhancement can support rolling upgrades where the new schema is backward-compatible, but this is explicitly out of scope for Stage 1. +- **ConfigMap as state store** — if the ConfigMap is accidentally deleted, the operator re-runs migration (which is safe — `openfga migrate` is idempotent). This is a feature, not a bug, but should be documented. diff --git a/docs/adr/003-declarative-store-lifecycle-crds.md b/docs/adr/003-declarative-store-lifecycle-crds.md new file mode 100644 index 00000000..a54ee44b --- /dev/null +++ b/docs/adr/003-declarative-store-lifecycle-crds.md @@ -0,0 +1,199 @@ +# ADR-003: Declarative Store Lifecycle Management via CRDs + +- **Status:** Proposed +- **Date:** 2026-04-06 +- **Deciders:** OpenFGA Helm Charts maintainers +- **Related ADR:** [ADR-001](001-adopt-openfga-operator.md) + +## Context + +OpenFGA is an authorization service. After deploying the server, teams must perform several runtime operations to make it usable: + +1. **Create a store** — a logical container for authorization data +2. **Write an authorization model** — the DSL that defines types, relations, and permissions +3. **Write tuples** — the relationship data that the model operates on (e.g., "user:anne is owner of document:budget") + +Today, these operations happen outside Kubernetes — through the OpenFGA API, CLI (`fga`), or custom scripts in CI pipelines. There is no declarative, Kubernetes-native way to manage them. + +This creates several problems: + +- **No GitOps for authorization config** — authorization models live in scripts or API calls, not in version-controlled manifests that ArgoCD/FluxCD sync. +- **No drift detection** — if someone modifies a model or tuple via the API, there's no controller to detect and reconcile the change. +- **No cross-team ownership** — each team that uses OpenFGA must build their own tooling to manage stores and models. There's no standard pattern. +- **Manual coordination** — deploying a new version of an application that needs a model change requires coordinating the Helm upgrade with a separate model push. + +### Alternatives Considered + +**A. CLI wrapper in CI pipelines** + +Use the `fga` CLI in a CI/CD step after `helm upgrade` to create stores, push models, and write tuples. + +*Pros:* No new Kubernetes components. Works with any CI system. +*Cons:* Imperative, not declarative. No drift detection. Each team builds their own pipeline. Model changes are not atomic with deployments. No visibility in Kubernetes tooling. + +**B. Helm post-install hook Job** + +Add a Helm hook Job that runs `fga` CLI commands after installation. + +*Pros:* Stays within the Helm ecosystem. +*Cons:* Helm hooks are the exact problem we're solving in ADR-002. Same ArgoCD/FluxCD incompatibilities. Hook Jobs are fire-and-forget with no reconciliation. + +**C. CRDs managed by the operator (selected)** + +Expose `FGAStore`, `FGAModel`, and `FGATuples` as Custom Resource Definitions. The operator watches these resources and reconciles them against the OpenFGA API. + +*Pros:* Fully declarative. GitOps-native. Continuous reconciliation. Standard Kubernetes patterns. Teams own their auth config as manifests. +*Cons:* Requires the operator (ADR-001). CRD design and reconciliation logic add development scope. Tuple reconciliation is complex. + +## Decision + +Introduce three CRDs, built in stages after the migration handling (ADR-002) is complete: + +### Stage 2: FGAStore + +```yaml +apiVersion: openfga.dev/v1alpha1 +kind: FGAStore +metadata: + name: my-app + namespace: my-team +spec: + # Reference to the OpenFGA instance + openfgaRef: + url: openfga.openfga-system.svc:8081 + credentialsRef: + name: openfga-api-credentials # Secret with API key or client credentials + # Store display name + name: "my-app-store" +status: + storeId: "01HXYZ..." + ready: true + conditions: + - type: Ready + status: "True" + lastTransitionTime: "2026-04-06T12:00:00Z" +``` + +**Controller behavior:** +- On create: call `CreateStore` API, store the returned store ID in `.status.storeId` +- On delete: call `DeleteStore` API (with finalizer to ensure cleanup) +- Idempotent: if a store with the same name exists, adopt it rather than creating a duplicate +- Status: set `Ready` condition when store is confirmed to exist + +### Stage 3: FGAModel + +```yaml +apiVersion: openfga.dev/v1alpha1 +kind: FGAModel +metadata: + name: my-app-model + namespace: my-team +spec: + storeRef: + name: my-app # References an FGAStore in the same namespace + model: | + model + schema 1.1 + type user + type organization + relations + define member: [user] + define admin: [user] + type document + relations + define reader: [user, organization#member] + define writer: [user, organization#admin] + define owner: [user] +status: + modelId: "01HABC..." + ready: true + lastWrittenHash: "sha256:a1b2c3..." # Hash of the model DSL to detect changes + conditions: + - type: Ready + status: "True" + - type: InSync + status: "True" +``` + +**Controller behavior:** +- On create/update: hash the model DSL. If hash differs from `.status.lastWrittenHash`, call `WriteAuthorizationModel` API +- Store the returned model ID in `.status.modelId` +- Model writes are append-only in OpenFGA (each write creates a new version), so this is safe +- Validation: optionally validate DSL syntax before calling the API (fail-fast with a clear error condition) +- The controller does NOT delete old model versions — OpenFGA retains model history + +### Stage 4: FGATuples + +```yaml +apiVersion: openfga.dev/v1alpha1 +kind: FGATuples +metadata: + name: my-app-base-tuples + namespace: my-team +spec: + storeRef: + name: my-app + tuples: + - user: "user:anne" + relation: "owner" + object: "document:budget" + - user: "team:engineering#member" + relation: "reader" + object: "folder:engineering-docs" + - user: "organization:acme#admin" + relation: "writer" + object: "folder:engineering-docs" +status: + writtenCount: 3 + ready: true + lastReconciled: "2026-04-06T12:00:00Z" + conditions: + - type: Ready + status: "True" + - type: InSync + status: "True" +``` + +**Controller behavior:** +- Maintain an **ownership model** — the controller tracks which tuples it wrote (via annotations or a status field). It only manages tuples it owns, never deleting tuples written by the application at runtime. +- On reconciliation: diff the desired tuples (from spec) against owned tuples in the store + - Tuples in spec but not in store → write them + - Tuples in store (owned) but not in spec → delete them + - Tuples in store but not owned → leave them alone +- Pagination: handle large tuple sets that exceed API response limits +- Batching: use `Write` API with batch operations to minimize API calls + +**Scope limitation:** `FGATuples` is intended for **base/static tuples** — organizational structure, role assignments, resource hierarchies. It is NOT intended to replace application-level tuple writes for dynamic data (e.g., per-request access grants). The ownership model ensures these two concerns don't interfere. + +### CRD Design Principles + +1. **Namespace-scoped** — all CRDs are namespaced, allowing teams to manage their own stores/models/tuples in their namespace +2. **Reference-based** — `FGAModel` and `FGATuples` reference an `FGAStore` by name, not by store ID. The controller resolves the reference. +3. **Status-driven** — controllers report state via `.status.conditions` following Kubernetes conventions (`Ready`, `InSync`, error conditions) +4. **Finalizers for cleanup** — `FGAStore` uses a finalizer to ensure the store is deleted from OpenFGA when the CR is deleted +5. **Idempotent** — all operations are safe to retry. Re-running reconciliation produces the same result. +6. **`v1alpha1` API version** — signals that the CRD schema may change. We will promote to `v1beta1` and `v1` as the design stabilizes. + +## Consequences + +### Positive + +- **GitOps-native authorization management** — stores, models, and tuples are Kubernetes resources that ArgoCD/FluxCD sync from Git +- **Drift detection and reconciliation** — the operator continuously ensures the actual state matches the declared state +- **Cross-team standardization** — every team uses the same CRDs, eliminating custom scripts and CI hacks +- **Atomic deployments** — a team can include `FGAModel` in their application's Helm chart; model updates deploy alongside code changes +- **Visibility** — `kubectl get fgastores`, `kubectl get fgamodels`, `kubectl describe fgatuples` provide instant visibility into authorization configuration +- **RBAC integration** — Kubernetes RBAC controls who can create/modify stores, models, and tuples per namespace + +### Negative + +- **Significant development scope** — three controllers, each with its own reconciliation logic, error handling, and tests +- **Tuple reconciliation complexity** — diffing and ownership tracking for tuples is the most complex piece; edge cases around partial failures, pagination, and large tuple sets +- **CRD upgrade burden** — CRD schema changes require careful migration; Helm does not upgrade CRDs automatically +- **API dependency** — the operator must be able to reach the OpenFGA API; network issues or API downtime affect reconciliation +- **Not suitable for all tuple management** — dynamic, application-driven tuples should still be written via the API, not CRDs. Users must understand this boundary. + +### Risks + +- **FGATuples at scale** — for stores with millions of tuples, the reconciliation diff could be expensive. The ownership model mitigates this (only diff owned tuples), but documentation must clearly state that `FGATuples` is for base/static data, not high-volume dynamic writes. +- **Multi-cluster** — if OpenFGA serves multiple clusters, CRDs in one cluster may conflict with CRDs in another pointing at the same store. This is out of scope for `v1alpha1` but should be considered for future versions. diff --git a/docs/adr/004-operator-deployment-model.md b/docs/adr/004-operator-deployment-model.md new file mode 100644 index 00000000..bedb12c1 --- /dev/null +++ b/docs/adr/004-operator-deployment-model.md @@ -0,0 +1,167 @@ +# ADR-004: Operator Deployment as Helm Subchart Dependency + +- **Status:** Proposed +- **Date:** 2026-04-06 +- **Deciders:** OpenFGA Helm Charts maintainers +- **Related ADR:** [ADR-001](001-adopt-openfga-operator.md) + +## Context + +The OpenFGA Operator (ADR-001) needs a deployment model — how do users install it alongside or independent of the OpenFGA server? + +There are several established patterns in the Kubernetes ecosystem: + +### Alternatives Considered + +**A. Standalone operator chart (install separately)** + +Users install the operator chart first, then install the OpenFGA chart. The operator watches for OpenFGA Deployments across namespaces. + +*Example:* +```bash +helm install openfga-operator openfga/openfga-operator -n openfga-system +helm install openfga openfga/openfga -n my-namespace +``` + +*Pros:* Clean separation of concerns. One operator instance serves multiple OpenFGA installations. Follows the OLM/OperatorHub pattern. +*Cons:* Two install steps. Ordering dependency — operator must exist before the chart is useful. Users must manage two releases. Harder to get started. + +**B. Operator bundled in the main chart (single chart, always installed)** + +The operator Deployment, RBAC, and CRDs are templates in the main OpenFGA chart. No subchart. + +*Pros:* Simplest for users — one chart, one install. No dependency management. +*Cons:* Chart becomes larger and harder to maintain. Users who manage the operator separately (e.g., cluster-wide) can't disable it. CRDs are tied to the application chart's release cycle. Multiple OpenFGA installations in the same cluster would deploy multiple operator instances. + +**C. Operator as a conditional subchart dependency (selected)** + +The operator is a separate Helm chart (`openfga-operator`) that the main chart declares as a conditional dependency. Enabled by default, but users can disable it. + +*Example:* +```bash +# Everything in one command +helm install openfga openfga/openfga \ + --set datastore.engine=postgres \ + --set operator.enabled=true + +# Or, operator managed separately +helm install openfga-operator openfga/openfga-operator -n openfga-system +helm install openfga openfga/openfga \ + --set operator.enabled=false +``` + +*Pros:* Single install for most users. Operator chart has its own versioning. Users can disable for standalone management. Clean separation in code. +*Cons:* Subchart dependency adds some Chart.yaml complexity. CRDs still need special handling (Helm's `crds/` directory or a pre-install hook). + +**D. OLM (Operator Lifecycle Manager) only** + +Publish the operator to OperatorHub. Users install via OLM. + +*Pros:* Standard pattern for OpenShift. Handles CRD upgrades, operator upgrades, and RBAC. +*Cons:* OLM is not available on all clusters (not standard on EKS, GKE, AKS). Adds a dependency on OLM itself. Doesn't help Helm-only users. + +## Decision + +The operator will be distributed as a **conditional Helm subchart dependency** of the main OpenFGA chart. + +### Chart Structure + +``` +helm-charts/ +├── charts/ +│ ├── openfga/ # Main chart (existing) +│ │ ├── Chart.yaml # Declares openfga-operator as dependency +│ │ ├── values.yaml # operator.enabled: true +│ │ ├── templates/ +│ │ └── crds/ # Empty in Stage 1 +│ │ +│ └── openfga-operator/ # Operator subchart (new) +│ ├── Chart.yaml +│ ├── values.yaml +│ ├── templates/ +│ │ ├── deployment.yaml +│ │ ├── serviceaccount.yaml +│ │ ├── clusterrole.yaml +│ │ └── clusterrolebinding.yaml +│ └── crds/ # CRDs added in Stages 2-4 +│ ├── fgastore.yaml +│ ├── fgamodel.yaml +│ └── fgatuples.yaml +``` + +### Dependency Declaration + +```yaml +# charts/openfga/Chart.yaml +dependencies: + - name: openfga-operator + version: "0.1.x" + repository: "oci://ghcr.io/openfga/helm-charts" + condition: operator.enabled +``` + +### CRD Handling + +Helm has specific behavior around CRDs: + +1. **`crds/` directory** — CRDs placed here are installed on `helm install` but are **never upgraded or deleted** by Helm. This is safe but requires manual CRD upgrades. + +2. **Pre-install/pre-upgrade hook Job** — a Job that runs `kubectl apply -f` on CRD manifests before the main install/upgrade. This handles upgrades but reintroduces Helm hooks (the problem ADR-002 solves). + +3. **Static manifests applied separately** — CRDs are published as a standalone YAML file. Users run `kubectl apply -f` before `helm install`. This is the pattern used by cert-manager, Istio, and Prometheus Operator. + +**Decision:** Use the `crds/` directory in the operator subchart for initial installation. Publish CRD manifests as a standalone artifact for upgrades. Document both paths clearly. + +```bash +# First install — Helm installs CRDs automatically +helm install openfga openfga/openfga + +# CRD upgrades — applied manually (Helm won't upgrade them) +kubectl apply -f https://github.com/openfga/helm-charts/releases/download/v0.2.0/crds.yaml +``` + +### Installation Modes + +| Mode | Command | Use case | +|------|---------|----------| +| **All-in-one** (default) | `helm install openfga openfga/openfga` | Most users. Single install, operator included. | +| **Operator disabled** | `helm install openfga openfga/openfga --set operator.enabled=false` | Operator managed separately or not needed (memory datastore). | +| **Operator standalone** | `helm install op openfga/openfga-operator -n openfga-system` | Cluster-wide operator serving multiple OpenFGA instances. | + +### Multi-Instance Considerations + +When multiple OpenFGA installations exist in the same cluster: + +- **All-in-one mode:** Each installation gets its own operator instance. The operator only watches resources in its own namespace. This is simple but wasteful. +- **Standalone mode:** One operator installation watches all namespaces (or a configured set). Individual OpenFGA installations set `operator.enabled=false`. This is more efficient for large clusters. + +The operator will support both modes via a `watchNamespace` configuration: + +```yaml +# Operator values +operator: + watchNamespace: "" # empty = watch own namespace only (all-in-one mode) + # watchNamespace: "" # or set to a specific namespace + # watchAllNamespaces: true # watch all namespaces (standalone mode) +``` + +## Consequences + +### Positive + +- **Single `helm install` for most users** — no ordering dependencies, no manual operator setup +- **Opt-out available** — `operator.enabled: false` for users who manage it separately or don't need it +- **Independent versioning** — operator chart has its own version; can be released on a different cadence than the main chart +- **Clean code separation** — operator code and templates are in their own chart directory +- **Standalone installation supported** — cluster admins can install one operator for multiple OpenFGA instances +- **Consistent with ecosystem** — this is the same pattern used by charts that depend on Bitnami PostgreSQL, Redis, etc. + +### Negative + +- **CRD upgrade complexity** — Helm does not upgrade CRDs; users must apply CRD manifests separately on operator upgrades +- **Multiple operators in all-in-one mode** — if a user installs OpenFGA in three namespaces, they get three operator pods (wasteful). Documentation should recommend standalone mode for multi-instance clusters. +- **Subchart value passing** — configuring the operator requires prefixed values (e.g., `openfga-operator.image.tag`), which is slightly less ergonomic than top-level values + +### Neutral + +- **OLM support is not excluded** — the operator can be published to OperatorHub in the future alongside the Helm distribution. The two are not mutually exclusive. diff --git a/docs/adr/README.md b/docs/adr/README.md new file mode 100644 index 00000000..298a9e32 --- /dev/null +++ b/docs/adr/README.md @@ -0,0 +1,180 @@ +# Architecture Decision Records + +This directory contains Architecture Decision Records (ADRs) for the OpenFGA Helm Charts project. + +ADRs are short documents that capture significant architectural decisions along with their context, alternatives considered, and consequences. They serve as a decision log — not a living design doc, but a point-in-time record of *why* a decision was made. + +We follow the format described by [Michael Nygard](https://cognitect.com/blog/2011/11/15/documenting-architecture-decisions). + +## Index + +| ADR | Title | Status | Date | +|-----|-------|--------|------| +| [ADR-001](001-adopt-openfga-operator.md) | Adopt a Kubernetes Operator for OpenFGA Lifecycle Management | Proposed | 2026-04-06 | +| [ADR-002](002-operator-managed-migrations.md) | Replace Helm Hook Migrations with Operator-Managed Migrations | Proposed | 2026-04-06 | +| [ADR-003](003-declarative-store-lifecycle-crds.md) | Declarative Store Lifecycle Management via CRDs | Proposed | 2026-04-06 | +| [ADR-004](004-operator-deployment-model.md) | Operator Deployment as Helm Subchart Dependency | Proposed | 2026-04-06 | + +--- + +## What is an ADR? + +An ADR captures a single architectural decision. It records: + +- **What** was decided +- **Why** it was decided (the context and constraints at the time) +- **What alternatives** were considered and why they were rejected +- **What consequences** follow from the decision (positive, negative, and neutral) + +ADRs are **immutable once accepted** — if a decision changes, you write a new ADR that supersedes the old one rather than editing it. This preserves the history of *why* things changed over time. + +## ADR Lifecycle + +``` +Proposed → Accepted → (optionally) Superseded or Deprecated + ↑ + │ feedback loop + │ + Discussion +``` + +### Statuses + +| Status | Meaning | +|--------|---------| +| **Proposed** | The ADR has been written and is open for discussion. No commitment has been made. | +| **Accepted** | The decision has been agreed upon by maintainers. Implementation can proceed. | +| **Deprecated** | The decision is no longer relevant (e.g., the feature was removed). | +| **Superseded by ADR-XXX** | A newer ADR has replaced this decision. The old ADR links to the new one. | + +## How to Propose an ADR + +1. **Create a branch** — e.g., `docs/adr-005-my-decision` + +2. **Copy the template** — use `000-template.md` as a starting point + +3. **Write the ADR** — fill in Context, Decision, and Consequences. Focus on *why*, not *how*. The most valuable part is the Alternatives Considered section — it shows reviewers what you evaluated and why you chose this path. + +4. **Assign a number** — use the next sequential number. Check the index above. + +5. **Open a pull request** — the PR is where discussion happens. Title it: `ADR-005: ` + +6. **Add to the index** — update the table in this README with the new entry (status: Proposed) + +### Proposing related ADRs together + +When multiple ADRs are part of a single cohesive proposal — e.g., a foundational decision and several downstream decisions that depend on it — they can be submitted in a single PR. This lets reviewers see the full picture instead of bouncing between separate PRs. + +When doing this: + +- **Explain the relationship in the PR description** — identify which ADR is the foundational decision and which are downstream. For example: "ADR-001 is the core decision to build an operator. ADR-002, 003, and 004 are downstream decisions about how the operator handles migrations, CRDs, and deployment." +- **Each ADR can be accepted or rejected independently** — a reviewer might approve the foundational decision but push back on a downstream one. If that happens, split the PR: merge the accepted ADRs and keep the contested ones open for further discussion. +- **Keep each ADR self-contained** — even though they're in the same PR, each ADR should stand on its own. A reader should be able to understand ADR-003 without reading ADR-002 first (though they may reference each other). + +## How to Give Feedback on an ADR + +ADR review happens in the **pull request**, not by editing the ADR directly. This keeps the discussion visible and linked to the decision. + +### As a reviewer + +- **Comment on the PR** — ask questions, challenge assumptions, suggest alternatives. Good review questions: + - "Did you consider X as an alternative?" + - "What happens if Y fails?" + - "This conflicts with how we do Z — can you address that?" + - "I agree with the decision but the consequence about X should mention Y" + +- **Request changes** if you believe the decision is wrong or incomplete + +- **Approve** when you're satisfied the decision is sound and well-documented + +### As the author responding to feedback + +- **Update the ADR in the PR** based on feedback: + - Add alternatives that reviewers suggested (with your evaluation of them) + - Expand the Consequences section if reviewers identified impacts you missed + - Clarify the Context if reviewers were confused about the problem + - Adjust the Decision if feedback reveals a better approach + +- **Do NOT delete feedback-driven changes** — if a reviewer raised a valid alternative and you addressed it, the ADR is stronger for including it + +- **Resolve PR comments** as you address them so reviewers can track progress + +### Reaching consensus + +- ADRs move to **Accepted** when maintainers approve the PR +- Not every maintainer needs to approve — follow the project's normal review standards +- If consensus can't be reached, escalate to a synchronous discussion (meeting, call) and record the outcome in the PR +- Disagreement is fine — document it in the Consequences section as a risk or trade-off rather than hiding it + +## How to Supersede an ADR + +When a decision needs to change: + +1. **Do NOT edit the original ADR** — it's a historical record + +2. **Write a new ADR** that references the old one: + ```markdown + - **Supersedes:** [ADR-002](002-operator-managed-migrations.md) + ``` + +3. **Update the old ADR's status** — change it to: + ```markdown + - **Status:** Superseded by [ADR-007](007-new-approach.md) + ``` + +4. **Update the index** in this README + +This way, anyone reading ADR-002 knows it's been replaced and can follow the link to understand what changed and why. + +## ADR Format + +Every ADR follows this structure: + +```markdown +# ADR-NNN: Title + +- **Status:** Proposed | Accepted | Deprecated | Superseded by ADR-XXX +- **Date:** YYYY-MM-DD +- **Deciders:** Who was involved in the decision +- **Related Issues:** GitHub issue references +- **Related ADR:** Links to related ADRs + +## Context + +What is the problem or situation that motivates this decision? +Include enough background that someone unfamiliar with the project +can understand why this decision matters. + +## Decision + +What is the decision and why was it chosen? + +### Alternatives Considered + +What other options were evaluated? Why were they rejected? +This is often the most valuable section — it prevents future +contributors from re-proposing rejected approaches. + +## Consequences + +### Positive +What improves as a result of this decision? + +### Negative +What gets harder or more complex? Be honest — every decision has costs. + +### Risks +What could go wrong? What assumptions might prove false? +``` + +## Template + +A blank template is available at [000-template.md](000-template.md). + +## Tips for Writing Good ADRs + +- **Keep it short** — an ADR is one decision, not a design doc. If it's longer than 2-3 pages, consider splitting it. +- **Focus on why, not how** — implementation details change; the reasoning behind the decision is what matters long-term. +- **Be honest about trade-offs** — an ADR that lists only positive consequences isn't credible. Every decision has costs. +- **Write for your future self** — in 18 months, you won't remember why you chose this. The ADR should tell you. +- **Not every decision needs an ADR** — use ADRs for decisions that are hard to reverse, affect multiple components, or where the reasoning isn't obvious from the code. From 93292f3c14acb3c93abe4bef1b7ff1498baef1cd Mon Sep 17 00:00:00 2001 From: Ed Milic <edmilic@gmail.com> Date: Fri, 10 Apr 2026 13:09:35 -0400 Subject: [PATCH 02/70] feat: add operator for migration orchestration (Stage 1) Replace Helm hook-based migrations with a lightweight Kubernetes operator that watches OpenFGA Deployments, detects version changes, and runs migrations as regular Jobs. - Go operator using controller-runtime (no CRDs) - Helm subchart with opt-in via operator.enabled (default false) - Dedicated migration ServiceAccount (separate from runtime) - Auto-recovery on database failure (delete/retry cycle) - GitHub Actions workflow for multi-arch image builds - Integration test values for local Kubernetes clusters Resolves #211, #107, #120, #100, #126 --- .github/workflows/operator.yml | 100 ++++++ charts/openfga-operator/Chart.yaml | 13 + charts/openfga-operator/crds/README.md | 4 + charts/openfga-operator/templates/NOTES.txt | 11 + .../openfga-operator/templates/_helpers.tpl | 72 ++++ .../templates/clusterrole.yaml | 31 ++ .../templates/clusterrolebinding.yaml | 14 + .../templates/deployment.yaml | 79 ++++ .../templates/serviceaccount.yaml | 13 + charts/openfga-operator/values.yaml | 59 +++ charts/openfga/Chart.lock | 7 +- charts/openfga/Chart.yaml | 4 + charts/openfga/templates/_helpers.tpl | 11 + charts/openfga/templates/deployment.yaml | 17 +- charts/openfga/templates/job.yaml | 2 +- charts/openfga/templates/rbac.yaml | 2 +- charts/openfga/templates/serviceaccount.yaml | 13 + charts/openfga/values.schema.json | 69 ++++ charts/openfga/values.yaml | 26 ++ operator/.dockerignore | 6 + operator/Dockerfile | 17 + operator/Makefile | 26 ++ operator/README.md | 129 +++++++ operator/cmd/main.go | 102 ++++++ operator/go.mod | 66 ++++ operator/go.sum | 171 +++++++++ operator/internal/controller/helpers.go | 269 ++++++++++++++ .../controller/migration_controller.go | 215 +++++++++++ .../controller/migration_controller_test.go | 336 ++++++++++++++++++ operator/tests/README.md | 186 ++++++++++ operator/tests/values-db-outage.yaml | 70 ++++ operator/tests/values-happy-path.yaml | 70 ++++ operator/tests/values-no-db.yaml | 23 ++ 33 files changed, 2225 insertions(+), 8 deletions(-) create mode 100644 .github/workflows/operator.yml create mode 100644 charts/openfga-operator/Chart.yaml create mode 100644 charts/openfga-operator/crds/README.md create mode 100644 charts/openfga-operator/templates/NOTES.txt create mode 100644 charts/openfga-operator/templates/_helpers.tpl create mode 100644 charts/openfga-operator/templates/clusterrole.yaml create mode 100644 charts/openfga-operator/templates/clusterrolebinding.yaml create mode 100644 charts/openfga-operator/templates/deployment.yaml create mode 100644 charts/openfga-operator/templates/serviceaccount.yaml create mode 100644 charts/openfga-operator/values.yaml create mode 100644 operator/.dockerignore create mode 100644 operator/Dockerfile create mode 100644 operator/Makefile create mode 100644 operator/README.md create mode 100644 operator/cmd/main.go create mode 100644 operator/go.mod create mode 100644 operator/go.sum create mode 100644 operator/internal/controller/helpers.go create mode 100644 operator/internal/controller/migration_controller.go create mode 100644 operator/internal/controller/migration_controller_test.go create mode 100644 operator/tests/README.md create mode 100644 operator/tests/values-db-outage.yaml create mode 100644 operator/tests/values-happy-path.yaml create mode 100644 operator/tests/values-no-db.yaml diff --git a/.github/workflows/operator.yml b/.github/workflows/operator.yml new file mode 100644 index 00000000..71f9e7f1 --- /dev/null +++ b/.github/workflows/operator.yml @@ -0,0 +1,100 @@ +name: Operator + +on: + push: + branches: + - main + paths: + - "operator/**" + pull_request: + paths: + - "operator/**" + workflow_dispatch: + inputs: + push_image: + description: "Push the operator image to GHCR" + type: boolean + default: true + +env: + IMAGE_NAME: ghcr.io/${{ github.repository_owner }}/openfga-operator + +jobs: + test: + runs-on: ubuntu-latest + permissions: + contents: read + steps: + - name: Checkout + uses: actions/checkout@v6 + + - name: Set up Go + uses: actions/setup-go@v5 + with: + go-version-file: operator/go.mod + cache-dependency-path: operator/go.sum + + - name: Run tests + working-directory: operator + run: go test ./... -v + + - name: Run vet + working-directory: operator + run: go vet ./... + + build-and-push: + needs: test + if: >- + (github.event_name == 'push' && github.ref == 'refs/heads/main') || + (github.event_name == 'workflow_dispatch' && inputs.push_image) + runs-on: ubuntu-latest + permissions: + contents: read + packages: write + steps: + - name: Checkout + uses: actions/checkout@v6 + + - name: Extract version from Chart.yaml + id: version + run: | + version=$(grep '^appVersion:' charts/openfga-operator/Chart.yaml | awk '{print $2}' | tr -d '"') + echo "version=${version}" >> "$GITHUB_OUTPUT" + short_sha="${GITHUB_SHA::7}" + echo "short_sha=${short_sha}" >> "$GITHUB_OUTPUT" + echo "Operator version: ${version} (sha: ${short_sha})" + + - name: Determine image tags + id: tags + run: | + if [[ "${{ github.ref }}" == "refs/heads/main" ]]; then + echo "tags=${{ env.IMAGE_NAME }}:${{ steps.version.outputs.version }},${{ env.IMAGE_NAME }}:latest" >> "$GITHUB_OUTPUT" + else + # Dev build — tag with version-sha to avoid clobbering release tags + echo "tags=${{ env.IMAGE_NAME }}:${{ steps.version.outputs.version }}-${{ steps.version.outputs.short_sha }}" >> "$GITHUB_OUTPUT" + fi + + - name: Set up Docker Buildx + uses: docker/setup-buildx-action@v3 + + - name: Login to GHCR + uses: docker/login-action@v4.1.0 + with: + registry: ghcr.io + username: ${{ github.actor }} + password: ${{ secrets.GITHUB_TOKEN }} + + - name: Build and push + uses: docker/build-push-action@v6 + with: + context: operator + push: true + platforms: linux/amd64,linux/arm64 + tags: ${{ steps.tags.outputs.tags }} + cache-from: type=gha + cache-to: type=gha,mode=max + labels: | + org.opencontainers.image.source=https://github.com/${{ github.repository }} + org.opencontainers.image.version=${{ steps.version.outputs.version }} + org.opencontainers.image.title=openfga-operator + org.opencontainers.image.description=OpenFGA Kubernetes operator for migration orchestration diff --git a/charts/openfga-operator/Chart.yaml b/charts/openfga-operator/Chart.yaml new file mode 100644 index 00000000..1bdacb03 --- /dev/null +++ b/charts/openfga-operator/Chart.yaml @@ -0,0 +1,13 @@ +apiVersion: v2 +name: openfga-operator +description: Helm chart for the OpenFGA Kubernetes operator. + +type: application +version: 0.1.0 +appVersion: "0.1.0" + +home: "https://openfga.github.io/helm-charts" +icon: https://github.com/openfga/community/raw/main/brand-assets/icon/color/openfga-icon-color.svg + +annotations: + artifacthub.io/license: Apache-2.0 diff --git a/charts/openfga-operator/crds/README.md b/charts/openfga-operator/crds/README.md new file mode 100644 index 00000000..060b0d0c --- /dev/null +++ b/charts/openfga-operator/crds/README.md @@ -0,0 +1,4 @@ +# CRDs + +This directory is reserved for Custom Resource Definitions added in later stages. +No CRDs are installed in Stage 1 (migration orchestration). diff --git a/charts/openfga-operator/templates/NOTES.txt b/charts/openfga-operator/templates/NOTES.txt new file mode 100644 index 00000000..8c398b1c --- /dev/null +++ b/charts/openfga-operator/templates/NOTES.txt @@ -0,0 +1,11 @@ +The openfga-operator has been deployed. + +NOTE: The operator container image ({{ .Values.image.repository }}:{{ .Values.image.tag | default .Chart.AppVersion }}) +does not exist yet. The operator pod will remain in ImagePullBackOff until +the Go binary is built and pushed. + +To check operator status: + kubectl get deployment --namespace {{ include "openfga-operator.namespace" . }} {{ include "openfga-operator.fullname" . }} + +To view operator logs (once the image is available): + kubectl logs --namespace {{ include "openfga-operator.namespace" . }} -l "app.kubernetes.io/name={{ include "openfga-operator.name" . }}" diff --git a/charts/openfga-operator/templates/_helpers.tpl b/charts/openfga-operator/templates/_helpers.tpl new file mode 100644 index 00000000..70d6e4c4 --- /dev/null +++ b/charts/openfga-operator/templates/_helpers.tpl @@ -0,0 +1,72 @@ +{{/* +Expand the name of the chart. +*/}} +{{- define "openfga-operator.name" -}} +{{- default .Chart.Name .Values.nameOverride | trunc 63 | trimSuffix "-" }} +{{- end }} + +{{/* +Create a default fully qualified app name. +We truncate at 63 chars because some Kubernetes name fields are limited to this (by the DNS naming spec). +If release name contains chart name it will be used as a full name. +*/}} +{{- define "openfga-operator.fullname" -}} +{{- if .Values.fullnameOverride }} +{{- .Values.fullnameOverride | trunc 63 | trimSuffix "-" }} +{{- else }} +{{- $name := default .Chart.Name .Values.nameOverride }} +{{- if contains $name .Release.Name }} +{{- .Release.Name | trunc 63 | trimSuffix "-" }} +{{- else }} +{{- printf "%s-%s" .Release.Name $name | trunc 63 | trimSuffix "-" }} +{{- end }} +{{- end }} +{{- end }} + +{{/* +Expand the namespace of the release. +Allows overriding it for multi-namespace deployments in combined charts. +*/}} +{{- define "openfga-operator.namespace" -}} +{{- default .Release.Namespace .Values.namespaceOverride | trunc 63 | trimSuffix "-" -}} +{{- end -}} + +{{/* +Create chart name and version as used by the chart label. +*/}} +{{- define "openfga-operator.chart" -}} +{{- printf "%s-%s" .Chart.Name .Chart.Version | replace "+" "_" | trunc 63 | trimSuffix "-" }} +{{- end }} + +{{/* +Common labels +*/}} +{{- define "openfga-operator.labels" -}} +helm.sh/chart: {{ include "openfga-operator.chart" . }} +{{ include "openfga-operator.selectorLabels" . }} +app.kubernetes.io/component: operator +{{- if .Chart.AppVersion }} +app.kubernetes.io/version: {{ .Chart.AppVersion | quote }} +{{- end }} +app.kubernetes.io/managed-by: {{ .Release.Service }} +app.kubernetes.io/part-of: openfga +{{- end }} + +{{/* +Selector labels +*/}} +{{- define "openfga-operator.selectorLabels" -}} +app.kubernetes.io/name: {{ include "openfga-operator.name" . }} +app.kubernetes.io/instance: {{ .Release.Name }} +{{- end }} + +{{/* +Create the name of the service account to use +*/}} +{{- define "openfga-operator.serviceAccountName" -}} +{{- if .Values.serviceAccount.create }} +{{- default (include "openfga-operator.fullname" .) .Values.serviceAccount.name }} +{{- else }} +{{- default "default" .Values.serviceAccount.name }} +{{- end }} +{{- end }} diff --git a/charts/openfga-operator/templates/clusterrole.yaml b/charts/openfga-operator/templates/clusterrole.yaml new file mode 100644 index 00000000..09d0fd7f --- /dev/null +++ b/charts/openfga-operator/templates/clusterrole.yaml @@ -0,0 +1,31 @@ +apiVersion: rbac.authorization.k8s.io/v1 +kind: ClusterRole +metadata: + name: {{ include "openfga-operator.fullname" . }} + labels: + {{- include "openfga-operator.labels" . | nindent 4 }} +rules: + - apiGroups: ["apps"] + resources: ["deployments"] + verbs: ["get", "list", "watch", "patch"] + - apiGroups: ["apps"] + resources: ["deployments/status"] + verbs: ["update"] + - apiGroups: ["batch"] + resources: ["jobs"] + verbs: ["get", "list", "watch", "create", "delete"] + - apiGroups: [""] + resources: ["configmaps"] + verbs: ["get", "list", "watch", "create", "update"] + - apiGroups: [""] + resources: ["secrets"] + verbs: ["get"] + - apiGroups: [""] + resources: ["serviceaccounts"] + verbs: ["get", "list", "create"] + - apiGroups: ["coordination.k8s.io"] + resources: ["leases"] + verbs: ["get", "list", "watch", "create", "update"] + - apiGroups: [""] + resources: ["events"] + verbs: ["create", "patch"] diff --git a/charts/openfga-operator/templates/clusterrolebinding.yaml b/charts/openfga-operator/templates/clusterrolebinding.yaml new file mode 100644 index 00000000..854521ab --- /dev/null +++ b/charts/openfga-operator/templates/clusterrolebinding.yaml @@ -0,0 +1,14 @@ +apiVersion: rbac.authorization.k8s.io/v1 +kind: ClusterRoleBinding +metadata: + name: {{ include "openfga-operator.fullname" . }} + labels: + {{- include "openfga-operator.labels" . | nindent 4 }} +roleRef: + apiGroup: rbac.authorization.k8s.io + kind: ClusterRole + name: {{ include "openfga-operator.fullname" . }} +subjects: + - kind: ServiceAccount + name: {{ include "openfga-operator.serviceAccountName" . }} + namespace: {{ include "openfga-operator.namespace" . }} diff --git a/charts/openfga-operator/templates/deployment.yaml b/charts/openfga-operator/templates/deployment.yaml new file mode 100644 index 00000000..ae8af0d5 --- /dev/null +++ b/charts/openfga-operator/templates/deployment.yaml @@ -0,0 +1,79 @@ +apiVersion: apps/v1 +kind: Deployment +metadata: + name: {{ include "openfga-operator.fullname" . }} + namespace: {{ include "openfga-operator.namespace" . }} + labels: + {{- include "openfga-operator.labels" . | nindent 4 }} +spec: + replicas: {{ .Values.replicaCount }} + selector: + matchLabels: + {{- include "openfga-operator.selectorLabels" . | nindent 6 }} + template: + metadata: + {{- with .Values.podAnnotations }} + annotations: + {{- toYaml . | nindent 8 }} + {{- end }} + labels: + {{- include "openfga-operator.labels" . | nindent 8 }} + spec: + {{- with .Values.imagePullSecrets }} + imagePullSecrets: + {{- toYaml . | nindent 8 }} + {{- end }} + serviceAccountName: {{ include "openfga-operator.serviceAccountName" . }} + {{- with .Values.podSecurityContext }} + securityContext: + {{- toYaml . | nindent 8 }} + {{- end }} + containers: + - name: operator + {{- with .Values.securityContext }} + securityContext: + {{- toYaml . | nindent 12 }} + {{- end }} + image: "{{ .Values.image.repository }}:{{ .Values.image.tag | default .Chart.AppVersion }}" + imagePullPolicy: {{ .Values.image.pullPolicy }} + args: + {{- if .Values.leaderElection.enabled }} + - --leader-elect + {{- end }} + {{- if .Values.watchNamespace }} + - --watch-namespace={{ .Values.watchNamespace }} + {{- else if .Values.watchAllNamespaces }} + - --watch-all-namespaces + {{- end }} + ports: + - name: healthz + containerPort: 8081 + protocol: TCP + livenessProbe: + httpGet: + path: /healthz + port: healthz + initialDelaySeconds: 15 + periodSeconds: 20 + readinessProbe: + httpGet: + path: /readyz + port: healthz + initialDelaySeconds: 5 + periodSeconds: 10 + {{- with .Values.resources }} + resources: + {{- toYaml . | nindent 12 }} + {{- end }} + {{- with .Values.nodeSelector }} + nodeSelector: + {{- toYaml . | nindent 8 }} + {{- end }} + {{- with .Values.affinity }} + affinity: + {{- toYaml . | nindent 8 }} + {{- end }} + {{- with .Values.tolerations }} + tolerations: + {{- toYaml . | nindent 8 }} + {{- end }} diff --git a/charts/openfga-operator/templates/serviceaccount.yaml b/charts/openfga-operator/templates/serviceaccount.yaml new file mode 100644 index 00000000..8b1f8941 --- /dev/null +++ b/charts/openfga-operator/templates/serviceaccount.yaml @@ -0,0 +1,13 @@ +{{- if .Values.serviceAccount.create -}} +apiVersion: v1 +kind: ServiceAccount +metadata: + name: {{ include "openfga-operator.serviceAccountName" . }} + namespace: {{ include "openfga-operator.namespace" . }} + labels: + {{- include "openfga-operator.labels" . | nindent 4 }} + {{- with .Values.serviceAccount.annotations }} + annotations: + {{- toYaml . | nindent 4 }} + {{- end }} +{{- end }} diff --git a/charts/openfga-operator/values.yaml b/charts/openfga-operator/values.yaml new file mode 100644 index 00000000..891ad574 --- /dev/null +++ b/charts/openfga-operator/values.yaml @@ -0,0 +1,59 @@ +replicaCount: 1 + +image: + repository: openfga/openfga-operator + pullPolicy: IfNotPresent + # -- Overrides the image tag whose default is the chart appVersion. + tag: "" + +imagePullSecrets: [] +nameOverride: "" +fullnameOverride: "" + +serviceAccount: + # -- Specifies whether a service account should be created. + create: true + # -- Annotations to add to the service account. + annotations: {} + # -- The name of the service account to use. + # If not set and create is true, a name is generated using the fullname template. + name: "" + +podAnnotations: {} + +podSecurityContext: {} + # runAsNonRoot: true + # seccompProfile: + # type: RuntimeDefault + +securityContext: {} + # capabilities: + # drop: + # - ALL + # readOnlyRootFilesystem: true + # runAsNonRoot: true + # runAsUser: 65532 + +# -- Constrain the operator to watch a single namespace. +# Leave empty to default to the release namespace. +watchNamespace: "" + +# -- Watch all namespaces. Overrides watchNamespace. +watchAllNamespaces: false + +leaderElection: + # -- Enable leader election for controller manager. + enabled: true + +resources: {} + # requests: + # cpu: 10m + # memory: 64Mi + # limits: + # memory: 128Mi + +nodeSelector: {} + +tolerations: [] + +affinity: {} diff --git a/charts/openfga/Chart.lock b/charts/openfga/Chart.lock index e82ffa5a..80114538 100644 --- a/charts/openfga/Chart.lock +++ b/charts/openfga/Chart.lock @@ -8,5 +8,8 @@ dependencies: - name: common repository: oci://registry-1.docker.io/bitnamicharts version: 2.13.3 -digest: sha256:4bbfb25821b0dfb6c70aabb5caf4c5ec7e6526261f93a8f531f507f1d4c43e3e -generated: "2026-03-18T11:41:40.1785546-04:00" +- name: openfga-operator + repository: file://../openfga-operator + version: 0.1.0 +digest: sha256:d502dc105790995a4368a049c0f593820d08f2f82dc9c9a70480a343c7affe8b +generated: "2026-04-10T11:45:16.638975-04:00" diff --git a/charts/openfga/Chart.yaml b/charts/openfga/Chart.yaml index f88f9ce5..1d5afc8c 100644 --- a/charts/openfga/Chart.yaml +++ b/charts/openfga/Chart.yaml @@ -29,3 +29,7 @@ dependencies: repository: oci://registry-1.docker.io/bitnamicharts tags: - bitnami-common + - name: openfga-operator + version: "0.1.0" + repository: "file://../openfga-operator" + condition: operator.enabled diff --git a/charts/openfga/templates/_helpers.tpl b/charts/openfga/templates/_helpers.tpl index 5889497a..35ad94a9 100644 --- a/charts/openfga/templates/_helpers.tpl +++ b/charts/openfga/templates/_helpers.tpl @@ -74,6 +74,17 @@ Create the name of the service account to use {{- end }} {{- end }} +{{/* +Create the name of the migration service account to use (operator mode only) +*/}} +{{- define "openfga.migrationServiceAccountName" -}} +{{- if .Values.migration.serviceAccount.name }} +{{- .Values.migration.serviceAccount.name | trunc 63 | trimSuffix "-" }} +{{- else }} +{{- printf "%s-migration" (include "openfga.fullname" .) | trunc 63 | trimSuffix "-" }} +{{- end }} +{{- end }} + {{/* Return true if a secret object should be created */}} diff --git a/charts/openfga/templates/deployment.yaml b/charts/openfga/templates/deployment.yaml index e6c1fff9..1cd20acf 100644 --- a/charts/openfga/templates/deployment.yaml +++ b/charts/openfga/templates/deployment.yaml @@ -4,12 +4,21 @@ metadata: name: {{ include "openfga.fullname" . }} labels: {{- include "openfga.labels" . | nindent 4 }} - {{- with .Values.annotations }} annotations: + {{- if .Values.operator.enabled }} + openfga.dev/desired-replicas: "{{ ternary 1 .Values.replicaCount (eq .Values.datastore.engine "memory") }}" + openfga.dev/migration-service-account: "{{ include "openfga.migrationServiceAccountName" . }}" + {{- end }} + {{- with .Values.annotations }} {{- toYaml . | nindent 4 }} - {{- end }} + {{- end }} spec: - {{- if not .Values.autoscaling.enabled }} + {{- if .Values.operator.enabled }} + {{- if .Values.autoscaling.enabled }} + {{- fail "operator.enabled and autoscaling.enabled cannot both be true" }} + {{- end }} + replicas: 0 + {{- else if not .Values.autoscaling.enabled }} replicas: {{ ternary 1 .Values.replicaCount (eq .Values.datastore.engine "memory")}} {{- end }} selector: @@ -37,7 +46,7 @@ spec: serviceAccountName: {{ include "openfga.serviceAccountName" . }} securityContext: {{- toYaml .Values.podSecurityContext | nindent 8 }} - {{ if or (and (has .Values.datastore.engine (list "postgres" "mysql")) .Values.datastore.applyMigrations .Values.datastore.waitForMigrations) .Values.extraInitContainers }} + {{ if and (not .Values.operator.enabled) (or (and (has .Values.datastore.engine (list "postgres" "mysql")) .Values.datastore.applyMigrations .Values.datastore.waitForMigrations) .Values.extraInitContainers) }} initContainers: {{- if and (has .Values.datastore.engine (list "postgres" "mysql")) .Values.datastore.applyMigrations .Values.datastore.waitForMigrations (eq .Values.datastore.migrationType "job") }} - name: wait-for-migration diff --git a/charts/openfga/templates/job.yaml b/charts/openfga/templates/job.yaml index fc70228d..d46d938f 100644 --- a/charts/openfga/templates/job.yaml +++ b/charts/openfga/templates/job.yaml @@ -1,4 +1,4 @@ -{{- if and (has .Values.datastore.engine (list "postgres" "mysql")) .Values.datastore.applyMigrations (eq .Values.datastore.migrationType "job") -}} +{{- if and (not .Values.operator.enabled) (has .Values.datastore.engine (list "postgres" "mysql")) .Values.datastore.applyMigrations (eq .Values.datastore.migrationType "job") -}} apiVersion: batch/v1 kind: Job metadata: diff --git a/charts/openfga/templates/rbac.yaml b/charts/openfga/templates/rbac.yaml index 3c8e0f8b..71d3c096 100644 --- a/charts/openfga/templates/rbac.yaml +++ b/charts/openfga/templates/rbac.yaml @@ -1,4 +1,4 @@ -{{- if .Values.serviceAccount.create -}} +{{- if and (not .Values.operator.enabled) .Values.serviceAccount.create -}} apiVersion: rbac.authorization.k8s.io/v1 kind: Role metadata: diff --git a/charts/openfga/templates/serviceaccount.yaml b/charts/openfga/templates/serviceaccount.yaml index bbe191c9..278be07b 100644 --- a/charts/openfga/templates/serviceaccount.yaml +++ b/charts/openfga/templates/serviceaccount.yaml @@ -10,3 +10,16 @@ metadata: {{- toYaml . | nindent 4 }} {{- end }} {{- end }} +{{- if and .Values.operator.enabled .Values.migration.serviceAccount.create }} +--- +apiVersion: v1 +kind: ServiceAccount +metadata: + name: {{ include "openfga.migrationServiceAccountName" . }} + labels: + {{- include "openfga.labels" . | nindent 4 }} + {{- with .Values.migration.serviceAccount.annotations }} + annotations: + {{- toYaml . | nindent 4 }} + {{- end }} +{{- end }} diff --git a/charts/openfga/values.schema.json b/charts/openfga/values.schema.json index 151cc21a..76b2cf9b 100644 --- a/charts/openfga/values.schema.json +++ b/charts/openfga/values.schema.json @@ -1289,6 +1289,75 @@ "type": "boolean", "description": "This value is not used by this chart, but allows a common pattern of enabling/disabling subchart dependencies (where OpenFGA is a subchart)", "default": false + }, + "operator": { + "type": "object", + "description": "Controls the openfga-operator subchart. When enabled, migration is managed by the operator instead of the Helm job hook.", + "properties": { + "enabled": { + "type": "boolean", + "description": "Enable the openfga-operator subchart for operator-managed migrations", + "default": false + } + } + }, + "openfga-operator": { + "type": "object", + "description": "Values passed through to the openfga-operator subchart" + }, + "migration": { + "type": "object", + "description": "Controls operator-driven migration behavior. Only used when operator.enabled is true.", + "properties": { + "enabled": { + "type": "boolean", + "description": "Enable operator-managed database migrations", + "default": true + }, + "timeout": { + "type": ["string", "null"], + "description": "Timeout passed to the migration Job as activeDeadlineSeconds", + "default": "" + }, + "backoffLimit": { + "type": "integer", + "description": "Number of retries before marking the migration as failed", + "default": 3 + }, + "ttlSecondsAfterFinished": { + "type": "integer", + "description": "Seconds to keep completed/failed migration Jobs before cleanup", + "default": 600 + }, + "serviceAccount": { + "type": "object", + "properties": { + "create": { + "type": "boolean", + "description": "Create a dedicated service account for migration Jobs", + "default": true + }, + "annotations": { + "type": "object", + "description": "Annotations to add to the migration service account", + "additionalProperties": { + "type": "string" + }, + "default": {} + }, + "name": { + "type": "string", + "description": "The name of the migration service account. Defaults to {fullname}-migration.", + "default": "" + } + } + }, + "resources": { + "type": "object", + "description": "Resource requests/limits for migration Job pods", + "default": {} + } + } } }, "additionalProperties": false diff --git a/charts/openfga/values.yaml b/charts/openfga/values.yaml index 75fa19d9..e843563a 100644 --- a/charts/openfga/values.yaml +++ b/charts/openfga/values.yaml @@ -385,6 +385,32 @@ testContainerSpec: {} # -- Array of extra K8s manifests to deploy ## Note: Supports use of custom Helm templates extraObjects: [] + +# -- operator controls the openfga-operator subchart. +# When enabled, migration is managed by the operator instead of the Helm job hook. +operator: + enabled: false + +# -- migration controls operator-driven migration behavior. +# Only used when operator.enabled is true. +migration: + enabled: true + # -- Timeout passed to the migration Job as activeDeadlineSeconds. + timeout: "" + # -- Number of retries before marking the migration as failed. + backoffLimit: 3 + # -- Seconds to keep completed/failed migration Jobs before cleanup. + ttlSecondsAfterFinished: 600 + serviceAccount: + # -- Create a dedicated service account for migration Jobs. + create: true + # -- Annotations to add to the migration service account. + annotations: {} + # -- The name of the migration service account. + # If not set and create is true, defaults to {fullname}-migration. + name: "" + # -- Resource requests/limits for migration Job pods. + resources: {} ## Example: Deploy a PostgreSQL instance for dev/test using official Docker images. ## For production, use a managed database service or an operator like CloudnativePG. ## Configure the chart to use the secret: diff --git a/operator/.dockerignore b/operator/.dockerignore new file mode 100644 index 00000000..3efb8a0e --- /dev/null +++ b/operator/.dockerignore @@ -0,0 +1,6 @@ +**/.git +**/.gitignore +**/README.md +**/LICENSE +**/Makefile +**/.dockerignore diff --git a/operator/Dockerfile b/operator/Dockerfile new file mode 100644 index 00000000..be9097eb --- /dev/null +++ b/operator/Dockerfile @@ -0,0 +1,17 @@ +FROM golang:1.25 AS builder + +WORKDIR /workspace +COPY go.mod go.sum ./ +RUN go mod download + +COPY cmd/ cmd/ +COPY internal/ internal/ + +RUN CGO_ENABLED=0 GOOS=linux go build -ldflags="-s -w" -o /operator ./cmd/ + +FROM gcr.io/distroless/static:nonroot +WORKDIR / +COPY --from=builder /operator . +USER 65532:65532 + +ENTRYPOINT ["/operator"] diff --git a/operator/Makefile b/operator/Makefile new file mode 100644 index 00000000..575bb7cd --- /dev/null +++ b/operator/Makefile @@ -0,0 +1,26 @@ +IMG ?= openfga/openfga-operator:dev + +.PHONY: build test vet fmt lint docker-build docker-push clean + +build: + go build -o bin/operator ./cmd/ + +test: + go test ./... -v + +vet: + go vet ./... + +fmt: + go fmt ./... + +lint: vet fmt + +docker-build: + docker build -t $(IMG) . + +docker-push: + docker push $(IMG) + +clean: + rm -rf bin/ diff --git a/operator/README.md b/operator/README.md new file mode 100644 index 00000000..a2b06996 --- /dev/null +++ b/operator/README.md @@ -0,0 +1,129 @@ +# OpenFGA Operator + +A Kubernetes operator that manages database migrations for OpenFGA deployments. Instead of relying on Helm hooks and init containers, the operator watches OpenFGA Deployments, detects version changes, and orchestrates migrations as regular Jobs. + +This is **Stage 1** of the operator — focused solely on migration orchestration. See [ADR-001](../docs/adr/001-adopt-operator.md) for the full roadmap. + +## How It Works + +1. The operator watches Deployments labeled `app.kubernetes.io/part-of: openfga` +2. When a version change is detected (comparing the container image tag to the `openfga-migration-status` ConfigMap), the operator: + - Keeps the Deployment at 0 replicas + - Creates a migration Job running `openfga migrate` + - Waits for the Job to complete + - Updates the ConfigMap with the new version + - Scales the Deployment up to the desired replica count +3. On failure, a `MigrationFailed` condition is set on the Deployment and replicas stay at 0 + +## Prerequisites + +- Go 1.25+ +- Docker +- Helm 3.6+ +- A Kubernetes cluster (Rancher Desktop, kind, etc.) + +## Development + +### Build + +```bash +cd operator +go build ./... +``` + +### Test + +```bash +go test ./... -v +``` + +### Lint + +```bash +go vet ./... +``` + +### Docker Image + +```bash +docker build -t openfga/openfga-operator:dev . +``` + +## Local Testing + +Integration test values and instructions are in [`tests/`](tests/). Three scenarios are provided: + +| Scenario | Values File | What It Tests | +|----------|-------------|---------------| +| Happy path | `tests/values-happy-path.yaml` | Full lifecycle: Postgres up, migration succeeds, OpenFGA scales to 3/3 | +| DB outage & recovery | `tests/values-db-outage.yaml` | Postgres starts at 0 replicas; scale it up later to verify self-healing | +| No database | `tests/values-no-db.yaml` | Permanent failure: operator retries without crashing, app stays at 0 | + +Quick start: + +```bash +# 1. Build the operator image +cd operator +docker build -t openfga/openfga-operator:dev . + +# 2. Update chart dependencies +cd .. +helm dependency update charts/openfga + +# 3. Run the happy-path test +kubectl create namespace openfga-test +helm install openfga-test charts/openfga -n openfga-test \ + -f operator/tests/values-happy-path.yaml + +# 4. Verify (wait ~30s) +kubectl get all -n openfga-test + +# 5. Clean up +helm uninstall openfga-test -n openfga-test +kubectl delete namespace openfga-test +``` + +See [`tests/README.md`](tests/README.md) for detailed verification steps and all three scenarios. + +## Project Structure + +``` +operator/ +├── cmd/ +│ └── main.go # Entry point, manager setup +├── internal/ +│ └── controller/ +│ ├── migration_controller.go # Reconciliation loop +│ ├── migration_controller_test.go # Unit tests +│ └── helpers.go # Job builder, scaling, ConfigMap helpers +├── Dockerfile # Multi-stage build (distroless runtime) +├── Makefile +├── go.mod +└── go.sum +``` + +## Configuration + +The operator accepts the following flags: + +| Flag | Default | Description | +|------|---------|-------------| +| `--leader-elect` | `false` | Enable leader election | +| `--watch-namespace` | `""` | Namespace to watch (defaults to release namespace) | +| `--watch-all-namespaces` | `false` | Watch all namespaces | +| `--metrics-bind-address` | `:8080` | Metrics endpoint address | +| `--health-probe-bind-address` | `:8081` | Health probe endpoint address | +| `--backoff-limit` | `3` | BackoffLimit for migration Jobs | +| `--active-deadline-seconds` | `300` | ActiveDeadlineSeconds for migration Jobs | +| `--ttl-seconds-after-finished` | `300` | TTLSecondsAfterFinished for migration Jobs | + +When deployed via the Helm subchart, these are configured through `values.yaml`. See `charts/openfga-operator/values.yaml` for all available options. + +## Annotations + +The operator reads these annotations from the OpenFGA Deployment: + +| Annotation | Description | +|------------|-------------| +| `openfga.dev/desired-replicas` | The replica count to restore after migration succeeds. Set by the Helm chart. | +| `openfga.dev/migration-service-account` | The ServiceAccount to use for migration Jobs. Defaults to the Deployment's SA. | diff --git a/operator/cmd/main.go b/operator/cmd/main.go new file mode 100644 index 00000000..e4c9d3ea --- /dev/null +++ b/operator/cmd/main.go @@ -0,0 +1,102 @@ +package main + +import ( + "flag" + "os" + + "k8s.io/apimachinery/pkg/runtime" + utilruntime "k8s.io/apimachinery/pkg/runtime/serializer" + clientgoscheme "k8s.io/client-go/kubernetes/scheme" + ctrl "sigs.k8s.io/controller-runtime" + "sigs.k8s.io/controller-runtime/pkg/cache" + "sigs.k8s.io/controller-runtime/pkg/healthz" + "sigs.k8s.io/controller-runtime/pkg/log/zap" + metricsserver "sigs.k8s.io/controller-runtime/pkg/metrics/server" + + "github.com/openfga/openfga-operator/internal/controller" +) + +var scheme = runtime.NewScheme() + +func init() { + _ = clientgoscheme.AddToScheme(scheme) + // Suppress unused import. + _ = utilruntime.CodecFactory{} +} + +func main() { + var ( + leaderElect bool + watchNamespace string + watchAllNamespaces bool + metricsAddr string + healthProbeAddr string + backoffLimit int + activeDeadline int + ttlAfterFinished int + ) + + flag.BoolVar(&leaderElect, "leader-elect", false, "Enable leader election for the controller manager.") + flag.StringVar(&watchNamespace, "watch-namespace", "", "Namespace to watch. Defaults to the release namespace.") + flag.BoolVar(&watchAllNamespaces, "watch-all-namespaces", false, "Watch all namespaces.") + flag.StringVar(&metricsAddr, "metrics-bind-address", ":8080", "The address the metric endpoint binds to.") + flag.StringVar(&healthProbeAddr, "health-probe-bind-address", ":8081", "The address the health probe endpoint binds to.") + flag.IntVar(&backoffLimit, "backoff-limit", int(controller.DefaultBackoffLimit), "BackoffLimit for migration Jobs.") + flag.IntVar(&activeDeadline, "active-deadline-seconds", int(controller.DefaultActiveDeadlineSeconds), "ActiveDeadlineSeconds for migration Jobs.") + flag.IntVar(&ttlAfterFinished, "ttl-seconds-after-finished", int(controller.DefaultTTLSecondsAfterFinished), "TTLSecondsAfterFinished for migration Jobs.") + + opts := zap.Options{Development: false} + opts.BindFlags(flag.CommandLine) + flag.Parse() + + ctrl.SetLogger(zap.New(zap.UseFlagOptions(&opts))) + logger := ctrl.Log.WithName("setup") + + // Configure cache namespace restrictions. + var cacheOpts cache.Options + if watchNamespace != "" && !watchAllNamespaces { + cacheOpts.DefaultNamespaces = map[string]cache.Config{ + watchNamespace: {}, + } + } + + mgr, err := ctrl.NewManager(ctrl.GetConfigOrDie(), ctrl.Options{ + Scheme: scheme, + Metrics: metricsserver.Options{BindAddress: metricsAddr}, + HealthProbeBindAddress: healthProbeAddr, + LeaderElection: leaderElect, + LeaderElectionID: "openfga-operator-leader", + Cache: cacheOpts, + }) + if err != nil { + logger.Error(err, "unable to create manager") + os.Exit(1) + } + + reconciler := &controller.MigrationReconciler{ + Client: mgr.GetClient(), + BackoffLimit: int32(backoffLimit), + ActiveDeadlineSeconds: int64(activeDeadline), + TTLSecondsAfterFinished: int32(ttlAfterFinished), + } + + if err := reconciler.SetupWithManager(mgr); err != nil { + logger.Error(err, "unable to create controller", "controller", "MigrationReconciler") + os.Exit(1) + } + + if err := mgr.AddHealthzCheck("healthz", healthz.Ping); err != nil { + logger.Error(err, "unable to set up health check") + os.Exit(1) + } + if err := mgr.AddReadyzCheck("readyz", healthz.Ping); err != nil { + logger.Error(err, "unable to set up readiness check") + os.Exit(1) + } + + logger.Info("starting manager") + if err := mgr.Start(ctrl.SetupSignalHandler()); err != nil { + logger.Error(err, "problem running manager") + os.Exit(1) + } +} diff --git a/operator/go.mod b/operator/go.mod new file mode 100644 index 00000000..fea96493 --- /dev/null +++ b/operator/go.mod @@ -0,0 +1,66 @@ +module github.com/openfga/openfga-operator + +go 1.25.6 + +require ( + k8s.io/api v0.35.3 + k8s.io/apimachinery v0.35.3 + k8s.io/client-go v0.35.3 + k8s.io/utils v0.0.0-20260319190234-28399d86e0b5 + sigs.k8s.io/controller-runtime v0.23.3 +) + +require ( + github.com/beorn7/perks v1.0.1 // indirect + github.com/cespare/xxhash/v2 v2.3.0 // indirect + github.com/davecgh/go-spew v1.1.1 // indirect + github.com/emicklei/go-restful/v3 v3.12.2 // indirect + github.com/evanphx/json-patch/v5 v5.9.11 // indirect + github.com/fsnotify/fsnotify v1.9.0 // indirect + github.com/fxamacker/cbor/v2 v2.9.0 // indirect + github.com/go-logr/logr v1.4.3 // indirect + github.com/go-logr/zapr v1.3.0 // indirect + github.com/go-openapi/jsonpointer v0.21.0 // indirect + github.com/go-openapi/jsonreference v0.20.2 // indirect + github.com/go-openapi/swag v0.23.0 // indirect + github.com/google/btree v1.1.3 // indirect + github.com/google/gnostic-models v0.7.0 // indirect + github.com/google/go-cmp v0.7.0 // indirect + github.com/google/uuid v1.6.0 // indirect + github.com/josharian/intern v1.0.0 // indirect + github.com/json-iterator/go v1.1.12 // indirect + github.com/mailru/easyjson v0.7.7 // indirect + github.com/modern-go/concurrent v0.0.0-20180306012644-bacd9c7ef1dd // indirect + github.com/modern-go/reflect2 v1.0.3-0.20250322232337-35a7c28c31ee // indirect + github.com/munnerz/goautoneg v0.0.0-20191010083416-a7dc8b61c822 // indirect + github.com/pmezard/go-difflib v1.0.0 // indirect + github.com/prometheus/client_golang v1.23.2 // indirect + github.com/prometheus/client_model v0.6.2 // indirect + github.com/prometheus/common v0.66.1 // indirect + github.com/prometheus/procfs v0.16.1 // indirect + github.com/spf13/pflag v1.0.9 // indirect + github.com/x448/float16 v0.8.4 // indirect + go.uber.org/multierr v1.11.0 // indirect + go.uber.org/zap v1.27.0 // indirect + go.yaml.in/yaml/v2 v2.4.3 // indirect + go.yaml.in/yaml/v3 v3.0.4 // indirect + golang.org/x/net v0.47.0 // indirect + golang.org/x/oauth2 v0.30.0 // indirect + golang.org/x/sync v0.18.0 // indirect + golang.org/x/sys v0.38.0 // indirect + golang.org/x/term v0.37.0 // indirect + golang.org/x/text v0.31.0 // indirect + golang.org/x/time v0.9.0 // indirect + gomodules.xyz/jsonpatch/v2 v2.4.0 // indirect + google.golang.org/protobuf v1.36.8 // indirect + gopkg.in/evanphx/json-patch.v4 v4.13.0 // indirect + gopkg.in/inf.v0 v0.9.1 // indirect + gopkg.in/yaml.v3 v3.0.1 // indirect + k8s.io/apiextensions-apiserver v0.35.0 // indirect + k8s.io/klog/v2 v2.130.1 // indirect + k8s.io/kube-openapi v0.0.0-20250910181357-589584f1c912 // indirect + sigs.k8s.io/json v0.0.0-20250730193827-2d320260d730 // indirect + sigs.k8s.io/randfill v1.0.0 // indirect + sigs.k8s.io/structured-merge-diff/v6 v6.3.2-0.20260122202528-d9cc6641c482 // indirect + sigs.k8s.io/yaml v1.6.0 // indirect +) diff --git a/operator/go.sum b/operator/go.sum new file mode 100644 index 00000000..79e74816 --- /dev/null +++ b/operator/go.sum @@ -0,0 +1,171 @@ +github.com/Masterminds/semver/v3 v3.4.0 h1:Zog+i5UMtVoCU8oKka5P7i9q9HgrJeGzI9SA1Xbatp0= +github.com/Masterminds/semver/v3 v3.4.0/go.mod h1:4V+yj/TJE1HU9XfppCwVMZq3I84lprf4nC11bSS5beM= +github.com/beorn7/perks v1.0.1 h1:VlbKKnNfV8bJzeqoa4cOKqO6bYr3WgKZxO8Z16+hsOM= +github.com/beorn7/perks v1.0.1/go.mod h1:G2ZrVWU2WbWT9wwq4/hrbKbnv/1ERSJQ0ibhJ6rlkpw= +github.com/cespare/xxhash/v2 v2.3.0 h1:UL815xU9SqsFlibzuggzjXhog7bL6oX9BbNZnL2UFvs= +github.com/cespare/xxhash/v2 v2.3.0/go.mod h1:VGX0DQ3Q6kWi7AoAeZDth3/j3BFtOZR5XLFGgcrjCOs= +github.com/creack/pty v1.1.9/go.mod h1:oKZEueFk5CKHvIhNR5MUki03XCEU+Q6VDXinZuGJ33E= +github.com/davecgh/go-spew v1.1.0/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSsI+c5H38= +github.com/davecgh/go-spew v1.1.1 h1:vj9j/u1bqnvCEfJOwUhtlOARqs3+rkHYY13jYWTU97c= +github.com/davecgh/go-spew v1.1.1/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSsI+c5H38= +github.com/emicklei/go-restful/v3 v3.12.2 h1:DhwDP0vY3k8ZzE0RunuJy8GhNpPL6zqLkDf9B/a0/xU= +github.com/emicklei/go-restful/v3 v3.12.2/go.mod h1:6n3XBCmQQb25CM2LCACGz8ukIrRry+4bhvbpWn3mrbc= +github.com/evanphx/json-patch v0.5.2 h1:xVCHIVMUu1wtM/VkR9jVZ45N3FhZfYMMYGorLCR8P3k= +github.com/evanphx/json-patch v0.5.2/go.mod h1:ZWS5hhDbVDyob71nXKNL0+PWn6ToqBHMikGIFbs31qQ= +github.com/evanphx/json-patch/v5 v5.9.11 h1:/8HVnzMq13/3x9TPvjG08wUGqBTmZBsCWzjTM0wiaDU= +github.com/evanphx/json-patch/v5 v5.9.11/go.mod h1:3j+LviiESTElxA4p3EMKAB9HXj3/XEtnUf6OZxqIQTM= +github.com/fsnotify/fsnotify v1.9.0 h1:2Ml+OJNzbYCTzsxtv8vKSFD9PbJjmhYF14k/jKC7S9k= +github.com/fsnotify/fsnotify v1.9.0/go.mod h1:8jBTzvmWwFyi3Pb8djgCCO5IBqzKJ/Jwo8TRcHyHii0= +github.com/fxamacker/cbor/v2 v2.9.0 h1:NpKPmjDBgUfBms6tr6JZkTHtfFGcMKsw3eGcmD/sapM= +github.com/fxamacker/cbor/v2 v2.9.0/go.mod h1:vM4b+DJCtHn+zz7h3FFp/hDAI9WNWCsZj23V5ytsSxQ= +github.com/go-logr/logr v1.4.3 h1:CjnDlHq8ikf6E492q6eKboGOC0T8CDaOvkHCIg8idEI= +github.com/go-logr/logr v1.4.3/go.mod h1:9T104GzyrTigFIr8wt5mBrctHMim0Nb2HLGrmQ40KvY= +github.com/go-logr/zapr v1.3.0 h1:XGdV8XW8zdwFiwOA2Dryh1gj2KRQyOOoNmBy4EplIcQ= +github.com/go-logr/zapr v1.3.0/go.mod h1:YKepepNBd1u/oyhd/yQmtjVXmm9uML4IXUgMOwR8/Gg= +github.com/go-openapi/jsonpointer v0.19.6/go.mod h1:osyAmYz/mB/C3I+WsTTSgw1ONzaLJoLCyoi6/zppojs= +github.com/go-openapi/jsonpointer v0.21.0 h1:YgdVicSA9vH5RiHs9TZW5oyafXZFc6+2Vc1rr/O9oNQ= +github.com/go-openapi/jsonpointer v0.21.0/go.mod h1:IUyH9l/+uyhIYQ/PXVA41Rexl+kOkAPDdXEYns6fzUY= +github.com/go-openapi/jsonreference v0.20.2 h1:3sVjiK66+uXK/6oQ8xgcRKcFgQ5KXa2KvnJRumpMGbE= +github.com/go-openapi/jsonreference v0.20.2/go.mod h1:Bl1zwGIM8/wsvqjsOQLJ/SH+En5Ap4rVB5KVcIDZG2k= +github.com/go-openapi/swag v0.22.3/go.mod h1:UzaqsxGiab7freDnrUUra0MwWfN/q7tE4j+VcZ0yl14= +github.com/go-openapi/swag v0.23.0 h1:vsEVJDUo2hPJ2tu0/Xc+4noaxyEffXNIs3cOULZ+GrE= +github.com/go-openapi/swag v0.23.0/go.mod h1:esZ8ITTYEsH1V2trKHjAN8Ai7xHb8RV+YSZ577vPjgQ= +github.com/go-task/slim-sprig/v3 v3.0.0 h1:sUs3vkvUymDpBKi3qH1YSqBQk9+9D/8M2mN1vB6EwHI= +github.com/go-task/slim-sprig/v3 v3.0.0/go.mod h1:W848ghGpv3Qj3dhTPRyJypKRiqCdHZiAzKg9hl15HA8= +github.com/google/btree v1.1.3 h1:CVpQJjYgC4VbzxeGVHfvZrv1ctoYCAI8vbl07Fcxlyg= +github.com/google/btree v1.1.3/go.mod h1:qOPhT0dTNdNzV6Z/lhRX0YXUafgPLFUh+gZMl761Gm4= +github.com/google/gnostic-models v0.7.0 h1:qwTtogB15McXDaNqTZdzPJRHvaVJlAl+HVQnLmJEJxo= +github.com/google/gnostic-models v0.7.0/go.mod h1:whL5G0m6dmc5cPxKc5bdKdEN3UjI7OUGxBlw57miDrQ= +github.com/google/go-cmp v0.7.0 h1:wk8382ETsv4JYUZwIsn6YpYiWiBsYLSJiTsyBybVuN8= +github.com/google/go-cmp v0.7.0/go.mod h1:pXiqmnSA92OHEEa9HXL2W4E7lf9JzCmGVUdgjX3N/iU= +github.com/google/gofuzz v1.0.0/go.mod h1:dBl0BpW6vV/+mYPU4Po3pmUjxk6FQPldtuIdl/M65Eg= +github.com/google/gofuzz v1.2.0 h1:xRy4A+RhZaiKjJ1bPfwQ8sedCA+YS2YcCHW6ec7JMi0= +github.com/google/gofuzz v1.2.0/go.mod h1:dBl0BpW6vV/+mYPU4Po3pmUjxk6FQPldtuIdl/M65Eg= +github.com/google/pprof v0.0.0-20250403155104-27863c87afa6 h1:BHT72Gu3keYf3ZEu2J0b1vyeLSOYI8bm5wbJM/8yDe8= +github.com/google/pprof v0.0.0-20250403155104-27863c87afa6/go.mod h1:boTsfXsheKC2y+lKOCMpSfarhxDeIzfZG1jqGcPl3cA= +github.com/google/uuid v1.6.0 h1:NIvaJDMOsjHA8n1jAhLSgzrAzy1Hgr+hNrb57e+94F0= +github.com/google/uuid v1.6.0/go.mod h1:TIyPZe4MgqvfeYDBFedMoGGpEw/LqOeaOT+nhxU+yHo= +github.com/josharian/intern v1.0.0 h1:vlS4z54oSdjm0bgjRigI+G1HpF+tI+9rE5LLzOg8HmY= +github.com/josharian/intern v1.0.0/go.mod h1:5DoeVV0s6jJacbCEi61lwdGj/aVlrQvzHFFd8Hwg//Y= +github.com/json-iterator/go v1.1.12 h1:PV8peI4a0ysnczrg+LtxykD8LfKY9ML6u2jnxaEnrnM= +github.com/json-iterator/go v1.1.12/go.mod h1:e30LSqwooZae/UwlEbR2852Gd8hjQvJoHmT4TnhNGBo= +github.com/klauspost/compress v1.18.0 h1:c/Cqfb0r+Yi+JtIEq73FWXVkRonBlf0CRNYc8Zttxdo= +github.com/klauspost/compress v1.18.0/go.mod h1:2Pp+KzxcywXVXMr50+X0Q/Lsb43OQHYWRCY2AiWywWQ= +github.com/kr/pretty v0.2.1/go.mod h1:ipq/a2n7PKx3OHsz4KJII5eveXtPO4qwEXGdVfWzfnI= +github.com/kr/pretty v0.3.1 h1:flRD4NNwYAUpkphVc1HcthR4KEIFJ65n8Mw5qdRn3LE= +github.com/kr/pretty v0.3.1/go.mod h1:hoEshYVHaxMs3cyo3Yncou5ZscifuDolrwPKZanG3xk= +github.com/kr/pty v1.1.1/go.mod h1:pFQYn66WHrOpPYNljwOMqo10TkYh1fy3cYio2l3bCsQ= +github.com/kr/text v0.1.0/go.mod h1:4Jbv+DJW3UT/LiOwJeYQe1efqtUx/iVham/4vfdArNI= +github.com/kr/text v0.2.0 h1:5Nx0Ya0ZqY2ygV366QzturHI13Jq95ApcVaJBhpS+AY= +github.com/kr/text v0.2.0/go.mod h1:eLer722TekiGuMkidMxC/pM04lWEeraHUUmBw8l2grE= +github.com/kylelemons/godebug v1.1.0 h1:RPNrshWIDI6G2gRW9EHilWtl7Z6Sb1BR0xunSBf0SNc= +github.com/kylelemons/godebug v1.1.0/go.mod h1:9/0rRGxNHcop5bhtWyNeEfOS8JIWk580+fNqagV/RAw= +github.com/mailru/easyjson v0.7.7 h1:UGYAvKxe3sBsEDzO8ZeWOSlIQfWFlxbzLZe7hwFURr0= +github.com/mailru/easyjson v0.7.7/go.mod h1:xzfreul335JAWq5oZzymOObrkdz5UnU4kGfJJLY9Nlc= +github.com/modern-go/concurrent v0.0.0-20180228061459-e0a39a4cb421/go.mod h1:6dJC0mAP4ikYIbvyc7fijjWJddQyLn8Ig3JB5CqoB9Q= +github.com/modern-go/concurrent v0.0.0-20180306012644-bacd9c7ef1dd h1:TRLaZ9cD/w8PVh93nsPXa1VrQ6jlwL5oN8l14QlcNfg= +github.com/modern-go/concurrent v0.0.0-20180306012644-bacd9c7ef1dd/go.mod h1:6dJC0mAP4ikYIbvyc7fijjWJddQyLn8Ig3JB5CqoB9Q= +github.com/modern-go/reflect2 v1.0.2/go.mod h1:yWuevngMOJpCy52FWWMvUC8ws7m/LJsjYzDa0/r8luk= +github.com/modern-go/reflect2 v1.0.3-0.20250322232337-35a7c28c31ee h1:W5t00kpgFdJifH4BDsTlE89Zl93FEloxaWZfGcifgq8= +github.com/modern-go/reflect2 v1.0.3-0.20250322232337-35a7c28c31ee/go.mod h1:yWuevngMOJpCy52FWWMvUC8ws7m/LJsjYzDa0/r8luk= +github.com/munnerz/goautoneg v0.0.0-20191010083416-a7dc8b61c822 h1:C3w9PqII01/Oq1c1nUAm88MOHcQC9l5mIlSMApZMrHA= +github.com/munnerz/goautoneg v0.0.0-20191010083416-a7dc8b61c822/go.mod h1:+n7T8mK8HuQTcFwEeznm/DIxMOiR9yIdICNftLE1DvQ= +github.com/onsi/ginkgo/v2 v2.27.2 h1:LzwLj0b89qtIy6SSASkzlNvX6WktqurSHwkk2ipF/Ns= +github.com/onsi/ginkgo/v2 v2.27.2/go.mod h1:ArE1D/XhNXBXCBkKOLkbsb2c81dQHCRcF5zwn/ykDRo= +github.com/onsi/gomega v1.38.2 h1:eZCjf2xjZAqe+LeWvKb5weQ+NcPwX84kqJ0cZNxok2A= +github.com/onsi/gomega v1.38.2/go.mod h1:W2MJcYxRGV63b418Ai34Ud0hEdTVXq9NW9+Sx6uXf3k= +github.com/pkg/errors v0.9.1 h1:FEBLx1zS214owpjy7qsBeixbURkuhQAwrK5UwLGTwt4= +github.com/pkg/errors v0.9.1/go.mod h1:bwawxfHBFNV+L2hUp1rHADufV3IMtnDRdf1r5NINEl0= +github.com/pmezard/go-difflib v1.0.0 h1:4DBwDE0NGyQoBHbLQYPwSUPoCMWR5BEzIk/f1lZbAQM= +github.com/pmezard/go-difflib v1.0.0/go.mod h1:iKH77koFhYxTK1pcRnkKkqfTogsbg7gZNVY4sRDYZ/4= +github.com/prometheus/client_golang v1.23.2 h1:Je96obch5RDVy3FDMndoUsjAhG5Edi49h0RJWRi/o0o= +github.com/prometheus/client_golang v1.23.2/go.mod h1:Tb1a6LWHB3/SPIzCoaDXI4I8UHKeFTEQ1YCr+0Gyqmg= +github.com/prometheus/client_model v0.6.2 h1:oBsgwpGs7iVziMvrGhE53c/GrLUsZdHnqNwqPLxwZyk= +github.com/prometheus/client_model v0.6.2/go.mod h1:y3m2F6Gdpfy6Ut/GBsUqTWZqCUvMVzSfMLjcu6wAwpE= +github.com/prometheus/common v0.66.1 h1:h5E0h5/Y8niHc5DlaLlWLArTQI7tMrsfQjHV+d9ZoGs= +github.com/prometheus/common v0.66.1/go.mod h1:gcaUsgf3KfRSwHY4dIMXLPV0K/Wg1oZ8+SbZk/HH/dA= +github.com/prometheus/procfs v0.16.1 h1:hZ15bTNuirocR6u0JZ6BAHHmwS1p8B4P6MRqxtzMyRg= +github.com/prometheus/procfs v0.16.1/go.mod h1:teAbpZRB1iIAJYREa1LsoWUXykVXA1KlTmWl8x/U+Is= +github.com/rogpeppe/go-internal v1.14.1 h1:UQB4HGPB6osV0SQTLymcB4TgvyWu6ZyliaW0tI/otEQ= +github.com/rogpeppe/go-internal v1.14.1/go.mod h1:MaRKkUm5W0goXpeCfT7UZI6fk/L7L7so1lCWt35ZSgc= +github.com/spf13/pflag v1.0.9 h1:9exaQaMOCwffKiiiYk6/BndUBv+iRViNW+4lEMi0PvY= +github.com/spf13/pflag v1.0.9/go.mod h1:McXfInJRrz4CZXVZOBLb0bTZqETkiAhM9Iw0y3An2Bg= +github.com/stretchr/objx v0.1.0/go.mod h1:HFkY916IF+rwdDfMAkV7OtwuqBVzrE8GR6GFx+wExME= +github.com/stretchr/objx v0.4.0/go.mod h1:YvHI0jy2hoMjB+UWwv71VJQ9isScKT/TqJzVSSt89Yw= +github.com/stretchr/objx v0.5.0/go.mod h1:Yh+to48EsGEfYuaHDzXPcE3xhTkx73EhmCGUpEOglKo= +github.com/stretchr/objx v0.5.2 h1:xuMeJ0Sdp5ZMRXx/aWO6RZxdr3beISkG5/G/aIRr3pY= +github.com/stretchr/objx v0.5.2/go.mod h1:FRsXN1f5AsAjCGJKqEizvkpNtU+EGNCLh3NxZ/8L+MA= +github.com/stretchr/testify v1.3.0/go.mod h1:M5WIy9Dh21IEIfnGCwXGc5bZfKNJtfHm1UVUgZn+9EI= +github.com/stretchr/testify v1.7.1/go.mod h1:6Fq8oRcR53rry900zMqJjRRixrwX3KX962/h/Wwjteg= +github.com/stretchr/testify v1.8.0/go.mod h1:yNjHg4UonilssWZ8iaSj1OCr/vHnekPRkoO+kdMU+MU= +github.com/stretchr/testify v1.8.1/go.mod h1:w2LPCIKwWwSfY2zedu0+kehJoqGctiVI29o6fzry7u4= +github.com/stretchr/testify v1.11.1 h1:7s2iGBzp5EwR7/aIZr8ao5+dra3wiQyKjjFuvgVKu7U= +github.com/stretchr/testify v1.11.1/go.mod h1:wZwfW3scLgRK+23gO65QZefKpKQRnfz6sD981Nm4B6U= +github.com/x448/float16 v0.8.4 h1:qLwI1I70+NjRFUR3zs1JPUCgaCXSh3SW62uAKT1mSBM= +github.com/x448/float16 v0.8.4/go.mod h1:14CWIYCyZA/cWjXOioeEpHeN/83MdbZDRQHoFcYsOfg= +go.uber.org/goleak v1.3.0 h1:2K3zAYmnTNqV73imy9J1T3WC+gmCePx2hEGkimedGto= +go.uber.org/goleak v1.3.0/go.mod h1:CoHD4mav9JJNrW/WLlf7HGZPjdw8EucARQHekz1X6bE= +go.uber.org/multierr v1.11.0 h1:blXXJkSxSSfBVBlC76pxqeO+LN3aDfLQo+309xJstO0= +go.uber.org/multierr v1.11.0/go.mod h1:20+QtiLqy0Nd6FdQB9TLXag12DsQkrbs3htMFfDN80Y= +go.uber.org/zap v1.27.0 h1:aJMhYGrd5QSmlpLMr2MftRKl7t8J8PTZPA732ud/XR8= +go.uber.org/zap v1.27.0/go.mod h1:GB2qFLM7cTU87MWRP2mPIjqfIDnGu+VIO4V/SdhGo2E= +go.yaml.in/yaml/v2 v2.4.3 h1:6gvOSjQoTB3vt1l+CU+tSyi/HOjfOjRLJ4YwYZGwRO0= +go.yaml.in/yaml/v2 v2.4.3/go.mod h1:zSxWcmIDjOzPXpjlTTbAsKokqkDNAVtZO0WOMiT90s8= +go.yaml.in/yaml/v3 v3.0.4 h1:tfq32ie2Jv2UxXFdLJdh3jXuOzWiL1fo0bu/FbuKpbc= +go.yaml.in/yaml/v3 v3.0.4/go.mod h1:DhzuOOF2ATzADvBadXxruRBLzYTpT36CKvDb3+aBEFg= +golang.org/x/mod v0.29.0 h1:HV8lRxZC4l2cr3Zq1LvtOsi/ThTgWnUk/y64QSs8GwA= +golang.org/x/mod v0.29.0/go.mod h1:NyhrlYXJ2H4eJiRy/WDBO6HMqZQ6q9nk4JzS3NuCK+w= +golang.org/x/net v0.47.0 h1:Mx+4dIFzqraBXUugkia1OOvlD6LemFo1ALMHjrXDOhY= +golang.org/x/net v0.47.0/go.mod h1:/jNxtkgq5yWUGYkaZGqo27cfGZ1c5Nen03aYrrKpVRU= +golang.org/x/oauth2 v0.30.0 h1:dnDm7JmhM45NNpd8FDDeLhK6FwqbOf4MLCM9zb1BOHI= +golang.org/x/oauth2 v0.30.0/go.mod h1:B++QgG3ZKulg6sRPGD/mqlHQs5rB3Ml9erfeDY7xKlU= +golang.org/x/sync v0.18.0 h1:kr88TuHDroi+UVf+0hZnirlk8o8T+4MrK6mr60WkH/I= +golang.org/x/sync v0.18.0/go.mod h1:9KTHXmSnoGruLpwFjVSX0lNNA75CykiMECbovNTZqGI= +golang.org/x/sys v0.38.0 h1:3yZWxaJjBmCWXqhN1qh02AkOnCQ1poK6oF+a7xWL6Gc= +golang.org/x/sys v0.38.0/go.mod h1:OgkHotnGiDImocRcuBABYBEXf8A9a87e/uXjp9XT3ks= +golang.org/x/term v0.37.0 h1:8EGAD0qCmHYZg6J17DvsMy9/wJ7/D/4pV/wfnld5lTU= +golang.org/x/term v0.37.0/go.mod h1:5pB4lxRNYYVZuTLmy8oR2BH8dflOR+IbTYFD8fi3254= +golang.org/x/text v0.31.0 h1:aC8ghyu4JhP8VojJ2lEHBnochRno1sgL6nEi9WGFGMM= +golang.org/x/text v0.31.0/go.mod h1:tKRAlv61yKIjGGHX/4tP1LTbc13YSec1pxVEWXzfoeM= +golang.org/x/time v0.9.0 h1:EsRrnYcQiGH+5FfbgvV4AP7qEZstoyrHB0DzarOQ4ZY= +golang.org/x/time v0.9.0/go.mod h1:3BpzKBy/shNhVucY/MWOyx10tF3SFh9QdLuxbVysPQM= +golang.org/x/tools v0.38.0 h1:Hx2Xv8hISq8Lm16jvBZ2VQf+RLmbd7wVUsALibYI/IQ= +golang.org/x/tools v0.38.0/go.mod h1:yEsQ/d/YK8cjh0L6rZlY8tgtlKiBNTL14pGDJPJpYQs= +gomodules.xyz/jsonpatch/v2 v2.4.0 h1:Ci3iUJyx9UeRx7CeFN8ARgGbkESwJK+KB9lLcWxY/Zw= +gomodules.xyz/jsonpatch/v2 v2.4.0/go.mod h1:AH3dM2RI6uoBZxn3LVrfvJ3E0/9dG4cSrbuBJT4moAY= +google.golang.org/protobuf v1.36.8 h1:xHScyCOEuuwZEc6UtSOvPbAT4zRh0xcNRYekJwfqyMc= +google.golang.org/protobuf v1.36.8/go.mod h1:fuxRtAxBytpl4zzqUh6/eyUujkJdNiuEkXntxiD/uRU= +gopkg.in/check.v1 v0.0.0-20161208181325-20d25e280405/go.mod h1:Co6ibVJAznAaIkqp8huTwlJQCZ016jof/cbN4VW5Yz0= +gopkg.in/check.v1 v1.0.0-20201130134442-10cb98267c6c h1:Hei/4ADfdWqJk1ZMxUNpqntNwaWcugrBjAiHlqqRiVk= +gopkg.in/check.v1 v1.0.0-20201130134442-10cb98267c6c/go.mod h1:JHkPIbrfpd72SG/EVd6muEfDQjcINNoR0C8j2r3qZ4Q= +gopkg.in/evanphx/json-patch.v4 v4.13.0 h1:czT3CmqEaQ1aanPc5SdlgQrrEIb8w/wwCvWWnfEbYzo= +gopkg.in/evanphx/json-patch.v4 v4.13.0/go.mod h1:p8EYWUEYMpynmqDbY58zCKCFZw8pRWMG4EsWvDvM72M= +gopkg.in/inf.v0 v0.9.1 h1:73M5CoZyi3ZLMOyDlQh031Cx6N9NDJ2Vvfl76EDAgDc= +gopkg.in/inf.v0 v0.9.1/go.mod h1:cWUDdTG/fYaXco+Dcufb5Vnc6Gp2YChqWtbxRZE0mXw= +gopkg.in/yaml.v3 v3.0.0-20200313102051-9f266ea9e77c/go.mod h1:K4uyk7z7BCEPqu6E+C64Yfv1cQ7kz7rIZviUmN+EgEM= +gopkg.in/yaml.v3 v3.0.1 h1:fxVm/GzAzEWqLHuvctI91KS9hhNmmWOoWu0XTYJS7CA= +gopkg.in/yaml.v3 v3.0.1/go.mod h1:K4uyk7z7BCEPqu6E+C64Yfv1cQ7kz7rIZviUmN+EgEM= +k8s.io/api v0.35.3 h1:pA2fiBc6+N9PDf7SAiluKGEBuScsTzd2uYBkA5RzNWQ= +k8s.io/api v0.35.3/go.mod h1:9Y9tkBcFwKNq2sxwZTQh1Njh9qHl81D0As56tu42GA4= +k8s.io/apiextensions-apiserver v0.35.0 h1:3xHk2rTOdWXXJM+RDQZJvdx0yEOgC0FgQ1PlJatA5T4= +k8s.io/apiextensions-apiserver v0.35.0/go.mod h1:E1Ahk9SADaLQ4qtzYFkwUqusXTcaV2uw3l14aqpL2LU= +k8s.io/apimachinery v0.35.3 h1:MeaUwQCV3tjKP4bcwWGgZ/cp/vpsRnQzqO6J6tJyoF8= +k8s.io/apimachinery v0.35.3/go.mod h1:jQCgFZFR1F4Ik7hvr2g84RTJSZegBc8yHgFWKn//hns= +k8s.io/client-go v0.35.3 h1:s1lZbpN4uI6IxeTM2cpdtrwHcSOBML1ODNTCCfsP1pg= +k8s.io/client-go v0.35.3/go.mod h1:RzoXkc0mzpWIDvBrRnD+VlfXP+lRzqQjCmKtiwZ8Q9c= +k8s.io/klog/v2 v2.130.1 h1:n9Xl7H1Xvksem4KFG4PYbdQCQxqc/tTUyrgXaOhHSzk= +k8s.io/klog/v2 v2.130.1/go.mod h1:3Jpz1GvMt720eyJH1ckRHK1EDfpxISzJ7I9OYgaDtPE= +k8s.io/kube-openapi v0.0.0-20250910181357-589584f1c912 h1:Y3gxNAuB0OBLImH611+UDZcmKS3g6CthxToOb37KgwE= +k8s.io/kube-openapi v0.0.0-20250910181357-589584f1c912/go.mod h1:kdmbQkyfwUagLfXIad1y2TdrjPFWp2Q89B3qkRwf/pQ= +k8s.io/utils v0.0.0-20260319190234-28399d86e0b5 h1:kBawHLSnx/mYHmRnNUf9d4CpjREbeZuxoSGOX/J+aYM= +k8s.io/utils v0.0.0-20260319190234-28399d86e0b5/go.mod h1:xDxuJ0whA3d0I4mf/C4ppKHxXynQ+fxnkmQH0vTHnuk= +sigs.k8s.io/controller-runtime v0.23.3 h1:VjB/vhoPoA9l1kEKZHBMnQF33tdCLQKJtydy4iqwZ80= +sigs.k8s.io/controller-runtime v0.23.3/go.mod h1:B6COOxKptp+YaUT5q4l6LqUJTRpizbgf9KSRNdQGns0= +sigs.k8s.io/json v0.0.0-20250730193827-2d320260d730 h1:IpInykpT6ceI+QxKBbEflcR5EXP7sU1kvOlxwZh5txg= +sigs.k8s.io/json v0.0.0-20250730193827-2d320260d730/go.mod h1:mdzfpAEoE6DHQEN0uh9ZbOCuHbLK5wOm7dK4ctXE9Tg= +sigs.k8s.io/randfill v1.0.0 h1:JfjMILfT8A6RbawdsK2JXGBR5AQVfd+9TbzrlneTyrU= +sigs.k8s.io/randfill v1.0.0/go.mod h1:XeLlZ/jmk4i1HRopwe7/aU3H5n1zNUcX6TM94b3QxOY= +sigs.k8s.io/structured-merge-diff/v6 v6.3.2-0.20260122202528-d9cc6641c482 h1:2WOzJpHUBVrrkDjU4KBT8n5LDcj824eX0I5UKcgeRUs= +sigs.k8s.io/structured-merge-diff/v6 v6.3.2-0.20260122202528-d9cc6641c482/go.mod h1:M3W8sfWvn2HhQDIbGWj3S099YozAsymCo/wrT5ohRUE= +sigs.k8s.io/yaml v1.6.0 h1:G8fkbMSAFqgEFgh4b1wmtzDnioxFCUgTZhlbj5P9QYs= +sigs.k8s.io/yaml v1.6.0/go.mod h1:796bPqUfzR/0jLAl6XjHl3Ck7MiyVv8dbTdyT3/pMf4= diff --git a/operator/internal/controller/helpers.go b/operator/internal/controller/helpers.go new file mode 100644 index 00000000..175805d0 --- /dev/null +++ b/operator/internal/controller/helpers.go @@ -0,0 +1,269 @@ +package controller + +import ( + "context" + "fmt" + "strconv" + "strings" + "time" + + appsv1 "k8s.io/api/apps/v1" + batchv1 "k8s.io/api/batch/v1" + corev1 "k8s.io/api/core/v1" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/utils/ptr" + "sigs.k8s.io/controller-runtime/pkg/client" +) + +const ( + // Labels used to discover OpenFGA Deployments. + LabelPartOf = "app.kubernetes.io/part-of" + LabelComponent = "app.kubernetes.io/component" + + LabelPartOfValue = "openfga" + LabelComponentValue = "authorization-controller" + + // Annotations set on the Deployment by the Helm chart / operator. + AnnotationDesiredReplicas = "openfga.dev/desired-replicas" + AnnotationMigrationServiceAccount = "openfga.dev/migration-service-account" + + // Defaults for migration Job configuration. + DefaultBackoffLimit int32 = 3 + DefaultActiveDeadlineSeconds int64 = 300 + DefaultTTLSecondsAfterFinished int32 = 300 +) + +// extractImageTag returns the tag portion of a container image reference. +// For "openfga/openfga:v1.14.0" it returns "v1.14.0". +// For "openfga/openfga@sha256:abc..." it returns the digest. +// If there is no tag or digest, it returns "latest". +func extractImageTag(image string) string { + // Handle digest references. + if idx := strings.LastIndex(image, "@"); idx != -1 { + return image[idx+1:] + } + + // Handle tag references — be careful not to split on the port in a registry URL. + // Find the last '/' to isolate the image name from the registry. + lastSlash := strings.LastIndex(image, "/") + nameAndTag := image + if lastSlash != -1 { + nameAndTag = image[lastSlash+1:] + } + + if idx := strings.LastIndex(nameAndTag, ":"); idx != -1 { + return nameAndTag[idx+1:] + } + + return "latest" +} + +// migrationConfigMapName returns the name of the ConfigMap used to track migration state. +func migrationConfigMapName(deploymentName string) string { + return deploymentName + "-migration-status" +} + +// migrationJobName returns the name of the migration Job. +func migrationJobName(deploymentName string) string { + return deploymentName + "-migrate" +} + +// buildMigrationJob constructs a migration Job for the given Deployment and version. +func buildMigrationJob( + deployment *appsv1.Deployment, + desiredVersion string, + backoffLimit int32, + activeDeadlineSeconds int64, + ttlSecondsAfterFinished int32, +) *batchv1.Job { + // Extract the main container's image and datastore env vars. + mainContainer := deployment.Spec.Template.Spec.Containers[0] + + // Determine the migration service account. + migrationSA := deployment.Annotations[AnnotationMigrationServiceAccount] + if migrationSA == "" { + migrationSA = deployment.Spec.Template.Spec.ServiceAccountName + } + + // Filter env vars — only pass datastore-related vars to the migration Job. + var datastoreEnvVars []corev1.EnvVar + for _, env := range mainContainer.Env { + if strings.HasPrefix(env.Name, "OPENFGA_DATASTORE_") { + datastoreEnvVars = append(datastoreEnvVars, env) + } + } + + return &batchv1.Job{ + ObjectMeta: metav1.ObjectMeta{ + Name: migrationJobName(deployment.Name), + Namespace: deployment.Namespace, + Labels: map[string]string{ + LabelPartOf: LabelPartOfValue, + LabelComponent: "migration", + "app.kubernetes.io/managed-by": "openfga-operator", + }, + OwnerReferences: []metav1.OwnerReference{ + { + APIVersion: "apps/v1", + Kind: "Deployment", + Name: deployment.Name, + UID: deployment.UID, + Controller: ptr.To(true), + BlockOwnerDeletion: ptr.To(true), + }, + }, + }, + Spec: batchv1.JobSpec{ + BackoffLimit: ptr.To(backoffLimit), + ActiveDeadlineSeconds: ptr.To(activeDeadlineSeconds), + TTLSecondsAfterFinished: ptr.To(ttlSecondsAfterFinished), + Template: corev1.PodTemplateSpec{ + ObjectMeta: metav1.ObjectMeta{ + Labels: map[string]string{ + LabelPartOf: LabelPartOfValue, + LabelComponent: "migration", + }, + }, + Spec: corev1.PodSpec{ + ServiceAccountName: migrationSA, + RestartPolicy: corev1.RestartPolicyNever, + Containers: []corev1.Container{ + { + Name: "migrate-database", + Image: mainContainer.Image, + Args: []string{"migrate"}, + Env: datastoreEnvVars, + }, + }, + // Inherit scheduling constraints from the parent Deployment. + NodeSelector: deployment.Spec.Template.Spec.NodeSelector, + Tolerations: deployment.Spec.Template.Spec.Tolerations, + Affinity: deployment.Spec.Template.Spec.Affinity, + }, + }, + }, + } +} + +// updateMigrationStatus creates or updates the migration-status ConfigMap. +func updateMigrationStatus( + ctx context.Context, + c client.Client, + deployment *appsv1.Deployment, + version string, + jobName string, +) error { + cmName := migrationConfigMapName(deployment.Name) + cm := &corev1.ConfigMap{ + ObjectMeta: metav1.ObjectMeta{ + Name: cmName, + Namespace: deployment.Namespace, + Labels: map[string]string{ + LabelPartOf: LabelPartOfValue, + LabelComponent: "migration", + "app.kubernetes.io/managed-by": "openfga-operator", + }, + OwnerReferences: []metav1.OwnerReference{ + { + APIVersion: "apps/v1", + Kind: "Deployment", + Name: deployment.Name, + UID: deployment.UID, + Controller: ptr.To(true), + BlockOwnerDeletion: ptr.To(true), + }, + }, + }, + Data: map[string]string{ + "version": version, + "migratedAt": time.Now().UTC().Format(time.RFC3339), + "jobName": jobName, + }, + } + + // Try to get existing ConfigMap first. + existing := &corev1.ConfigMap{} + err := c.Get(ctx, client.ObjectKeyFromObject(cm), existing) + if err != nil { + if client.IgnoreNotFound(err) != nil { + return fmt.Errorf("getting migration status ConfigMap: %w", err) + } + // ConfigMap doesn't exist — create it. + if createErr := c.Create(ctx, cm); createErr != nil { + return fmt.Errorf("creating migration status ConfigMap: %w", createErr) + } + return nil + } + + // Update existing ConfigMap. + existing.Data = cm.Data + existing.Labels = cm.Labels + if updateErr := c.Update(ctx, existing); updateErr != nil { + return fmt.Errorf("updating migration status ConfigMap: %w", updateErr) + } + return nil +} + +// ensureDeploymentScaled ensures the Deployment is scaled to the desired replica count. +// The desired count is read from the AnnotationDesiredReplicas annotation. +// Returns true if the Deployment was already at the desired scale. +func ensureDeploymentScaled(ctx context.Context, c client.Client, deployment *appsv1.Deployment) (bool, error) { + desiredStr, ok := deployment.Annotations[AnnotationDesiredReplicas] + if !ok || desiredStr == "" { + // No annotation — nothing to do. The Deployment may not have been scaled down yet. + return true, nil + } + + desired, err := strconv.ParseInt(desiredStr, 10, 32) + if err != nil { + return false, fmt.Errorf("parsing desired replicas annotation: %w", err) + } + + desiredInt32 := int32(desired) + current := int32(1) + if deployment.Spec.Replicas != nil { + current = *deployment.Spec.Replicas + } + + if current == desiredInt32 { + return true, nil + } + + patch := client.MergeFrom(deployment.DeepCopy()) + deployment.Spec.Replicas = ptr.To(desiredInt32) + if patchErr := c.Patch(ctx, deployment, patch); patchErr != nil { + return false, fmt.Errorf("scaling deployment to %d replicas: %w", desiredInt32, patchErr) + } + return false, nil +} + +// scaleDeploymentToZero scales the Deployment to 0 replicas, storing the current +// desired count in an annotation so it can be restored later. +func scaleDeploymentToZero(ctx context.Context, c client.Client, deployment *appsv1.Deployment) error { + if deployment.Spec.Replicas != nil && *deployment.Spec.Replicas == 0 { + return nil // Already at zero. + } + + patch := client.MergeFrom(deployment.DeepCopy()) + + // Store the current desired replica count before zeroing. + currentReplicas := int32(1) + if deployment.Spec.Replicas != nil { + currentReplicas = *deployment.Spec.Replicas + } + + // Only store if not already stored (avoid overwriting with 0 on re-reconciliation). + if _, ok := deployment.Annotations[AnnotationDesiredReplicas]; !ok { + if deployment.Annotations == nil { + deployment.Annotations = make(map[string]string) + } + deployment.Annotations[AnnotationDesiredReplicas] = strconv.FormatInt(int64(currentReplicas), 10) + } + + deployment.Spec.Replicas = ptr.To(int32(0)) + + if err := c.Patch(ctx, deployment, patch); err != nil { + return fmt.Errorf("scaling deployment to 0: %w", err) + } + return nil +} diff --git a/operator/internal/controller/migration_controller.go b/operator/internal/controller/migration_controller.go new file mode 100644 index 00000000..935e4ffc --- /dev/null +++ b/operator/internal/controller/migration_controller.go @@ -0,0 +1,215 @@ +package controller + +import ( + "context" + "fmt" + "time" + + appsv1 "k8s.io/api/apps/v1" + batchv1 "k8s.io/api/batch/v1" + corev1 "k8s.io/api/core/v1" + apierrors "k8s.io/apimachinery/pkg/api/errors" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/apimachinery/pkg/types" + ctrl "sigs.k8s.io/controller-runtime" + "sigs.k8s.io/controller-runtime/pkg/builder" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/handler" + "sigs.k8s.io/controller-runtime/pkg/log" + "sigs.k8s.io/controller-runtime/pkg/predicate" + "sigs.k8s.io/controller-runtime/pkg/reconcile" +) + +// MigrationReconciler watches OpenFGA Deployments and orchestrates database +// migrations when the application version changes. +type MigrationReconciler struct { + client.Client + + // BackoffLimit for migration Jobs. + BackoffLimit int32 + // ActiveDeadlineSeconds for migration Jobs. + ActiveDeadlineSeconds int64 + // TTLSecondsAfterFinished for migration Jobs. + TTLSecondsAfterFinished int32 +} + +// Reconcile handles a single reconciliation for an OpenFGA Deployment. +func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) (ctrl.Result, error) { + logger := log.FromContext(ctx) + + // 1. Get the OpenFGA Deployment. + deployment := &appsv1.Deployment{} + if err := r.Get(ctx, req.NamespacedName, deployment); err != nil { + if apierrors.IsNotFound(err) { + return ctrl.Result{}, nil + } + return ctrl.Result{}, err + } + + // 2. Extract the desired version from the Deployment's image tag. + if len(deployment.Spec.Template.Spec.Containers) == 0 { + logger.Info("deployment has no containers, skipping") + return ctrl.Result{}, nil + } + desiredVersion := extractImageTag(deployment.Spec.Template.Spec.Containers[0].Image) + + // 3. Check current migration status from ConfigMap. + configMap := &corev1.ConfigMap{} + cmName := migrationConfigMapName(req.Name) + err := r.Get(ctx, types.NamespacedName{Name: cmName, Namespace: req.Namespace}, configMap) + + currentVersion := "" + if err == nil { + currentVersion = configMap.Data["version"] + } else if !apierrors.IsNotFound(err) { + return ctrl.Result{}, fmt.Errorf("getting migration status: %w", err) + } + + // 4. If versions match, ensure Deployment is scaled up and return. + if currentVersion == desiredVersion { + logger.V(1).Info("migration up to date", "version", desiredVersion) + if _, scaleErr := ensureDeploymentScaled(ctx, r.Client, deployment); scaleErr != nil { + return ctrl.Result{}, scaleErr + } + return ctrl.Result{}, nil + } + + logger.Info("migration needed", "currentVersion", currentVersion, "desiredVersion", desiredVersion) + + // 5. Ensure the Deployment is scaled to zero before migrating. + if err := scaleDeploymentToZero(ctx, r.Client, deployment); err != nil { + return ctrl.Result{}, err + } + + // 6. Check if a migration Job already exists. + jobName := migrationJobName(req.Name) + job := &batchv1.Job{} + err = r.Get(ctx, types.NamespacedName{Name: jobName, Namespace: req.Namespace}, job) + + if apierrors.IsNotFound(err) { + // Create the migration Job. + job = buildMigrationJob( + deployment, + desiredVersion, + r.BackoffLimit, + r.ActiveDeadlineSeconds, + r.TTLSecondsAfterFinished, + ) + if createErr := r.Create(ctx, job); createErr != nil { + return ctrl.Result{}, fmt.Errorf("creating migration job: %w", createErr) + } + logger.Info("created migration job", "job", jobName, "version", desiredVersion) + return ctrl.Result{RequeueAfter: 5 * time.Second}, nil + } else if err != nil { + return ctrl.Result{}, fmt.Errorf("getting migration job: %w", err) + } + + // 7. Check Job status. + if job.Status.Succeeded >= 1 { + logger.Info("migration succeeded", "version", desiredVersion) + + // Update migration status ConfigMap. + if statusErr := updateMigrationStatus(ctx, r.Client, deployment, desiredVersion, jobName); statusErr != nil { + return ctrl.Result{}, statusErr + } + + // Scale Deployment back up. + if _, scaleErr := ensureDeploymentScaled(ctx, r.Client, deployment); scaleErr != nil { + return ctrl.Result{}, scaleErr + } + + return ctrl.Result{}, nil + } + + backoffLimit := r.BackoffLimit + if job.Spec.BackoffLimit != nil { + backoffLimit = *job.Spec.BackoffLimit + } + + if job.Status.Failed >= backoffLimit { + logger.Error(nil, "migration job failed, will delete and retry", "job", jobName, "version", desiredVersion) + + // Set condition so kubectl describe shows the failure. + setMigrationFailedCondition(deployment, desiredVersion) + if patchErr := r.Status().Update(ctx, deployment); patchErr != nil { + logger.Error(patchErr, "failed to set MigrationFailed condition") + } + + // Delete the failed Job so a fresh one is created on the next reconcile. + // This allows auto-recovery when the database comes back. + propagation := metav1.DeletePropagationBackground + if delErr := r.Delete(ctx, job, &client.DeleteOptions{ + PropagationPolicy: &propagation, + }); delErr != nil && !apierrors.IsNotFound(delErr) { + return ctrl.Result{}, fmt.Errorf("deleting failed migration job: %w", delErr) + } + logger.Info("deleted failed migration job, will retry", "job", jobName) + + // Requeue after a longer delay to avoid tight retry loops. + return ctrl.Result{RequeueAfter: 60 * time.Second}, nil + } + + // 8. Job still running — requeue. + logger.V(1).Info("migration job in progress", "job", jobName) + return ctrl.Result{RequeueAfter: 10 * time.Second}, nil +} + +// setMigrationFailedCondition sets a MigrationFailed condition on the Deployment. +func setMigrationFailedCondition(deployment *appsv1.Deployment, version string) { + condition := appsv1.DeploymentCondition{ + Type: "MigrationFailed", + Status: corev1.ConditionTrue, + LastTransitionTime: metav1.Now(), + Reason: "MigrationJobFailed", + Message: fmt.Sprintf("Database migration failed for version %s. Check migration job logs.", version), + } + + // Replace existing MigrationFailed condition if present. + for i, c := range deployment.Status.Conditions { + if c.Type == "MigrationFailed" { + deployment.Status.Conditions[i] = condition + return + } + } + deployment.Status.Conditions = append(deployment.Status.Conditions, condition) +} + +// SetupWithManager sets up the controller with the Manager. +func (r *MigrationReconciler) SetupWithManager(mgr ctrl.Manager) error { + // Only watch Deployments that are part of OpenFGA. + labelPredicate, err := predicate.LabelSelectorPredicate(metav1.LabelSelector{ + MatchLabels: map[string]string{ + LabelPartOf: LabelPartOfValue, + LabelComponent: LabelComponentValue, + }, + }) + if err != nil { + return fmt.Errorf("creating label predicate: %w", err) + } + + return ctrl.NewControllerManagedBy(mgr). + For(&appsv1.Deployment{}, builder.WithPredicates(labelPredicate)). + Owns(&batchv1.Job{}). + Watches(&corev1.ConfigMap{}, handler.EnqueueRequestsFromMapFunc( + func(ctx context.Context, obj client.Object) []reconcile.Request { + // Only watch ConfigMaps that are migration status ConfigMaps. + if obj.GetLabels()[LabelPartOf] != LabelPartOfValue || + obj.GetLabels()["app.kubernetes.io/managed-by"] != "openfga-operator" { + return nil + } + // Map back to the owning Deployment. + for _, ref := range obj.GetOwnerReferences() { + if ref.Kind == "Deployment" { + return []reconcile.Request{ + {NamespacedName: types.NamespacedName{ + Name: ref.Name, + Namespace: obj.GetNamespace(), + }}, + } + } + } + return nil + }, + )). + Complete(r) +} diff --git a/operator/internal/controller/migration_controller_test.go b/operator/internal/controller/migration_controller_test.go new file mode 100644 index 00000000..dc0f0c56 --- /dev/null +++ b/operator/internal/controller/migration_controller_test.go @@ -0,0 +1,336 @@ +package controller + +import ( + "context" + "testing" + "time" + + appsv1 "k8s.io/api/apps/v1" + batchv1 "k8s.io/api/batch/v1" + corev1 "k8s.io/api/core/v1" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/apimachinery/pkg/runtime" + "k8s.io/apimachinery/pkg/types" + "k8s.io/utils/ptr" + clientgoscheme "k8s.io/client-go/kubernetes/scheme" + ctrl "sigs.k8s.io/controller-runtime" + "sigs.k8s.io/controller-runtime/pkg/client/fake" +) + +func newScheme() *runtime.Scheme { + s := runtime.NewScheme() + _ = clientgoscheme.AddToScheme(s) + return s +} + +func newTestDeployment(name, namespace, image string, replicas int32) *appsv1.Deployment { + return &appsv1.Deployment{ + ObjectMeta: metav1.ObjectMeta{ + Name: name, + Namespace: namespace, + UID: "test-uid-123", + Labels: map[string]string{ + LabelPartOf: LabelPartOfValue, + LabelComponent: LabelComponentValue, + }, + Annotations: map[string]string{}, + }, + Spec: appsv1.DeploymentSpec{ + Replicas: ptr.To(replicas), + Selector: &metav1.LabelSelector{ + MatchLabels: map[string]string{"app": "openfga"}, + }, + Template: corev1.PodTemplateSpec{ + ObjectMeta: metav1.ObjectMeta{ + Labels: map[string]string{"app": "openfga"}, + }, + Spec: corev1.PodSpec{ + ServiceAccountName: "openfga", + Containers: []corev1.Container{ + { + Name: "openfga", + Image: image, + Env: []corev1.EnvVar{ + {Name: "OPENFGA_DATASTORE_ENGINE", Value: "postgres"}, + {Name: "OPENFGA_DATASTORE_URI", Value: "postgres://localhost/openfga"}, + {Name: "OPENFGA_LOG_LEVEL", Value: "info"}, + }, + }, + }, + }, + }, + }, + } +} + +func newReconciler(objects ...runtime.Object) *MigrationReconciler { + scheme := newScheme() + clientBuilder := fake.NewClientBuilder().WithScheme(scheme) + for _, obj := range objects { + clientBuilder = clientBuilder.WithRuntimeObjects(obj) + } + return &MigrationReconciler{ + Client: clientBuilder.Build(), + BackoffLimit: DefaultBackoffLimit, + ActiveDeadlineSeconds: DefaultActiveDeadlineSeconds, + TTLSecondsAfterFinished: DefaultTTLSecondsAfterFinished, + } +} + +func TestReconcile_FirstInstall_CreatesJob(t *testing.T) { + // Given: a Deployment with no migration-status ConfigMap. + dep := newTestDeployment("openfga", "default", "openfga/openfga:v1.14.0", 0) + r := newReconciler(dep) + + // When: reconciling. + result, err := r.Reconcile(context.Background(), ctrl.Request{ + NamespacedName: types.NamespacedName{Name: "openfga", Namespace: "default"}, + }) + + // Then: a migration Job should be created and requeue requested. + if err != nil { + t.Fatalf("unexpected error: %v", err) + } + if result.RequeueAfter == 0 { + t.Error("expected requeue, got none") + } + + // Verify the Job was created. + job := &batchv1.Job{} + if err := r.Get(context.Background(), types.NamespacedName{ + Name: "openfga-migrate", Namespace: "default", + }, job); err != nil { + t.Fatalf("expected migration job to be created: %v", err) + } + + if job.Spec.Template.Spec.Containers[0].Image != "openfga/openfga:v1.14.0" { + t.Errorf("expected job image openfga/openfga:v1.14.0, got %s", job.Spec.Template.Spec.Containers[0].Image) + } + + if job.Spec.Template.Spec.Containers[0].Args[0] != "migrate" { + t.Errorf("expected job args [migrate], got %v", job.Spec.Template.Spec.Containers[0].Args) + } + + // Verify only datastore env vars were passed. + for _, env := range job.Spec.Template.Spec.Containers[0].Env { + if env.Name == "OPENFGA_LOG_LEVEL" { + t.Error("non-datastore env var OPENFGA_LOG_LEVEL should not be passed to migration job") + } + } +} + +func TestReconcile_VersionMatch_ScalesUp(t *testing.T) { + // Given: a Deployment at 0 replicas with matching migration-status ConfigMap. + dep := newTestDeployment("openfga", "default", "openfga/openfga:v1.14.0", 0) + dep.Annotations[AnnotationDesiredReplicas] = "3" + + cm := &corev1.ConfigMap{ + ObjectMeta: metav1.ObjectMeta{ + Name: "openfga-migration-status", + Namespace: "default", + }, + Data: map[string]string{ + "version": "v1.14.0", + "migratedAt": "2026-04-06T12:00:00Z", + "jobName": "openfga-migrate", + }, + } + + r := newReconciler(dep, cm) + + // When: reconciling. + result, err := r.Reconcile(context.Background(), ctrl.Request{ + NamespacedName: types.NamespacedName{Name: "openfga", Namespace: "default"}, + }) + + // Then: no error, no requeue. + if err != nil { + t.Fatalf("unexpected error: %v", err) + } + if result.RequeueAfter != 0 { + t.Error("expected no requeue when versions match") + } + + // Verify Deployment was scaled up. + updated := &appsv1.Deployment{} + if err := r.Get(context.Background(), types.NamespacedName{ + Name: "openfga", Namespace: "default", + }, updated); err != nil { + t.Fatalf("getting deployment: %v", err) + } + if *updated.Spec.Replicas != 3 { + t.Errorf("expected 3 replicas, got %d", *updated.Spec.Replicas) + } +} + +func TestReconcile_JobSucceeded_UpdatesConfigMapAndScalesUp(t *testing.T) { + // Given: a Deployment at 0 replicas, no ConfigMap, and a succeeded migration Job. + dep := newTestDeployment("openfga", "default", "openfga/openfga:v1.14.0", 0) + dep.Annotations[AnnotationDesiredReplicas] = "3" + + job := &batchv1.Job{ + ObjectMeta: metav1.ObjectMeta{ + Name: "openfga-migrate", + Namespace: "default", + OwnerReferences: []metav1.OwnerReference{ + { + APIVersion: "apps/v1", + Kind: "Deployment", + Name: "openfga", + UID: "test-uid-123", + }, + }, + }, + Spec: batchv1.JobSpec{ + BackoffLimit: ptr.To(int32(3)), + Template: corev1.PodTemplateSpec{ + Spec: corev1.PodSpec{ + Containers: []corev1.Container{{Name: "migrate", Image: "openfga/openfga:v1.14.0"}}, + RestartPolicy: corev1.RestartPolicyNever, + }, + }, + }, + Status: batchv1.JobStatus{ + Succeeded: 1, + }, + } + + r := newReconciler(dep, job) + + // When: reconciling. + _, err := r.Reconcile(context.Background(), ctrl.Request{ + NamespacedName: types.NamespacedName{Name: "openfga", Namespace: "default"}, + }) + + // Then: no error. + if err != nil { + t.Fatalf("unexpected error: %v", err) + } + + // Verify ConfigMap was created. + cm := &corev1.ConfigMap{} + if err := r.Get(context.Background(), types.NamespacedName{ + Name: "openfga-migration-status", Namespace: "default", + }, cm); err != nil { + t.Fatalf("expected ConfigMap to be created: %v", err) + } + if cm.Data["version"] != "v1.14.0" { + t.Errorf("expected version v1.14.0 in ConfigMap, got %s", cm.Data["version"]) + } + + // Verify Deployment was scaled up. + updated := &appsv1.Deployment{} + if err := r.Get(context.Background(), types.NamespacedName{ + Name: "openfga", Namespace: "default", + }, updated); err != nil { + t.Fatalf("getting deployment: %v", err) + } + if *updated.Spec.Replicas != 3 { + t.Errorf("expected 3 replicas, got %d", *updated.Spec.Replicas) + } +} + +func TestReconcile_JobFailed_DeletesJobAndRequeues(t *testing.T) { + // Given: a Deployment at 0 replicas and a failed migration Job. + dep := newTestDeployment("openfga", "default", "openfga/openfga:v1.14.0", 0) + dep.Annotations[AnnotationDesiredReplicas] = "3" + + job := &batchv1.Job{ + ObjectMeta: metav1.ObjectMeta{ + Name: "openfga-migrate", + Namespace: "default", + OwnerReferences: []metav1.OwnerReference{ + { + APIVersion: "apps/v1", + Kind: "Deployment", + Name: "openfga", + UID: "test-uid-123", + }, + }, + }, + Spec: batchv1.JobSpec{ + BackoffLimit: ptr.To(int32(3)), + Template: corev1.PodTemplateSpec{ + Spec: corev1.PodSpec{ + Containers: []corev1.Container{{Name: "migrate", Image: "openfga/openfga:v1.14.0"}}, + RestartPolicy: corev1.RestartPolicyNever, + }, + }, + }, + Status: batchv1.JobStatus{ + Failed: 3, + }, + } + + r := newReconciler(dep, job) + + // When: reconciling. + result, err := r.Reconcile(context.Background(), ctrl.Request{ + NamespacedName: types.NamespacedName{Name: "openfga", Namespace: "default"}, + }) + + // Then: no error, but requeue after 60s for retry. + if err != nil { + t.Fatalf("unexpected error: %v", err) + } + if result.RequeueAfter != 60*time.Second { + t.Errorf("expected 60s requeue, got %v", result.RequeueAfter) + } + + // Verify Deployment was NOT scaled up — still at 0. + updated := &appsv1.Deployment{} + if getErr := r.Get(context.Background(), types.NamespacedName{ + Name: "openfga", Namespace: "default", + }, updated); getErr != nil { + t.Fatalf("getting deployment: %v", getErr) + } + if *updated.Spec.Replicas != 0 { + t.Errorf("expected 0 replicas after failed migration, got %d", *updated.Spec.Replicas) + } + + // Verify the failed Job was deleted. + deletedJob := &batchv1.Job{} + if getErr := r.Get(context.Background(), types.NamespacedName{ + Name: "openfga-migrate", Namespace: "default", + }, deletedJob); getErr == nil { + t.Error("expected failed migration job to be deleted") + } +} + +func TestReconcile_DeploymentNotFound_NoError(t *testing.T) { + r := newReconciler() + + result, err := r.Reconcile(context.Background(), ctrl.Request{ + NamespacedName: types.NamespacedName{Name: "nonexistent", Namespace: "default"}, + }) + + if err != nil { + t.Fatalf("unexpected error: %v", err) + } + if result.RequeueAfter != 0 { + t.Error("expected no requeue for missing deployment") + } +} + +func TestExtractImageTag(t *testing.T) { + tests := []struct { + image string + expected string + }{ + {"openfga/openfga:v1.14.0", "v1.14.0"}, + {"openfga/openfga:latest", "latest"}, + {"openfga/openfga", "latest"}, + {"ghcr.io/openfga/openfga:v1.14.0", "v1.14.0"}, + {"registry.example.com:5000/openfga/openfga:v1.14.0", "v1.14.0"}, + {"openfga/openfga@sha256:abcdef1234567890", "sha256:abcdef1234567890"}, + } + + for _, tt := range tests { + t.Run(tt.image, func(t *testing.T) { + got := extractImageTag(tt.image) + if got != tt.expected { + t.Errorf("extractImageTag(%q) = %q, want %q", tt.image, got, tt.expected) + } + }) + } +} diff --git a/operator/tests/README.md b/operator/tests/README.md new file mode 100644 index 00000000..e4377a99 --- /dev/null +++ b/operator/tests/README.md @@ -0,0 +1,186 @@ +# Local Integration Tests + +Manual integration tests for the OpenFGA operator on a local Kubernetes cluster (Rancher Desktop, kind, minikube, etc.). + +## Prerequisites + +- A running local Kubernetes cluster +- Helm 3.6+ +- The operator image built locally: + ```bash + cd operator + docker build -t openfga/openfga-operator:dev . + ``` +- Chart dependencies updated: + ```bash + helm dependency update charts/openfga + ``` + +All test values files use `imagePullPolicy: Never`, so the locally-built image must be available to the cluster's container runtime. On Rancher Desktop (dockerd) and Docker Desktop this works automatically. For kind, load the image first: + +```bash +kind load docker-image openfga/openfga-operator:dev +``` + +## Test Scenarios + +### 1. Happy Path + +Deploys OpenFGA with a Postgres instance. The operator should run the migration and scale OpenFGA up within ~30 seconds. + +```bash +kubectl create namespace openfga-test +helm install openfga-test charts/openfga -n openfga-test \ + -f operator/tests/values-happy-path.yaml +``` + +**Expected outcome:** + +| Resource | State | +|----------|-------| +| `openfga-test-openfga-operator` | `1/1 Running` | +| `openfga-test-postgres` | `1/1 Running` | +| `openfga-test-migrate-xxxxx` | `0/1 Completed` | +| `openfga-test` (OpenFGA) | `3/3 Running` | + +**Verify:** + +```bash +# All resources healthy +kubectl get all -n openfga-test + +# Operator logs show full lifecycle +kubectl logs -n openfga-test deployment/openfga-test-openfga-operator + +# Migration status recorded +kubectl get configmap openfga-test-migration-status -n openfga-test -o jsonpath='{.data}' + +# Database tables created +kubectl exec -n openfga-test deployment/openfga-test-postgres -- \ + psql -U openfga -d openfga -c '\dt' + +# OpenFGA responding +kubectl run curl-test --image=curlimages/curl -n openfga-test \ + --rm -it --restart=Never -- curl -s http://openfga-test:8080/healthz +# Expected: {"status":"SERVING"} +``` + +**Clean up:** + +```bash +helm uninstall openfga-test -n openfga-test +kubectl delete namespace openfga-test +``` + +--- + +### 2. Database Outage and Recovery + +Deploys OpenFGA with a Postgres instance scaled to 0 replicas (simulating a database that isn't ready yet). The operator should retry migrations until Postgres becomes available, then self-heal. + +```bash +kubectl create namespace openfga-test +helm install openfga-test charts/openfga -n openfga-test \ + -f operator/tests/values-db-outage.yaml +``` + +**Expected behavior while Postgres is down:** + +- Migration Job runs and fails (each pod times out after ~60s) +- After 3 failures (backoffLimit), the operator: + - Sets `MigrationFailed: True` condition on the Deployment + - Deletes the failed Job + - Creates a fresh Job after a 60-second delay +- This cycle repeats indefinitely +- OpenFGA stays at 0 replicas throughout (safe — no unmigrated app running) + +**Watch the failure cycle:** + +```bash +# Check deployment conditions +kubectl get deployment openfga-test -n openfga-test \ + -o jsonpath='{range .status.conditions[*]}{.type}: {.status} - {.message}{"\n"}{end}' + +# Watch operator logs for delete/retry cycle +kubectl logs -n openfga-test deployment/openfga-test-openfga-operator -f +# Look for: +# "migration job failed, will delete and retry" +# "deleted failed migration job, will retry" +# "created migration job" +``` + +**Bring Postgres back (after a few minutes):** + +```bash +kubectl scale deployment openfga-test-postgres -n openfga-test --replicas=1 +``` + +**Expected recovery (within ~60s of Postgres becoming ready):** + +- The currently running migration pod connects and succeeds +- Operator updates the ConfigMap with the new version +- Operator scales OpenFGA to 3/3 replicas +- `{"status":"SERVING"}` from the health endpoint + +**Verify recovery:** + +```bash +# OpenFGA should be 3/3 Running +kubectl get all -n openfga-test + +# Migration status recorded +kubectl get configmap openfga-test-migration-status -n openfga-test -o jsonpath='{.data}' + +# Health check +kubectl run curl-test --image=curlimages/curl -n openfga-test \ + --rm -it --restart=Never -- curl -s http://openfga-test:8080/healthz +``` + +**Clean up:** + +```bash +helm uninstall openfga-test -n openfga-test +kubectl delete namespace openfga-test +``` + +--- + +### 3. No Database (Permanent Failure) + +Deploys OpenFGA pointing at a Postgres hostname that doesn't exist. The operator should continuously retry without crashing or leaving the app in a broken state. + +```bash +kubectl create namespace openfga-test +helm install openfga-test charts/openfga -n openfga-test \ + -f operator/tests/values-no-db.yaml +``` + +**Expected behavior:** + +- Migration Jobs fail repeatedly (DNS resolution fails for `postgres-does-not-exist`) +- Operator sets `MigrationFailed: True` on the Deployment +- Operator deletes failed Jobs and retries every ~60 seconds +- OpenFGA stays at 0 replicas indefinitely — never starts against an unmigrated database + +This scenario verifies the operator doesn't crash-loop or consume excessive resources when the database is permanently unavailable. + +**Verify:** + +```bash +# OpenFGA at 0/0, operator at 1/1 +kubectl get deployments -n openfga-test + +# MigrationFailed condition present +kubectl get deployment openfga-test -n openfga-test \ + -o jsonpath='{range .status.conditions[*]}{.type}: {.status} - {.message}{"\n"}{end}' + +# Operator logs show retry cycle +kubectl logs -n openfga-test deployment/openfga-test-openfga-operator --tail=20 +``` + +**Clean up:** + +```bash +helm uninstall openfga-test -n openfga-test +kubectl delete namespace openfga-test +``` diff --git a/operator/tests/values-db-outage.yaml b/operator/tests/values-db-outage.yaml new file mode 100644 index 00000000..a7c59720 --- /dev/null +++ b/operator/tests/values-db-outage.yaml @@ -0,0 +1,70 @@ +# Test values: Postgres deployed but scaled to 0 (simulates DB outage) +operator: + enabled: true + +openfga-operator: + image: + repository: openfga/openfga-operator + tag: dev + pullPolicy: Never + resources: + requests: + cpu: 10m + memory: 64Mi + +datastore: + engine: postgres + uri: "postgres://openfga:changeme@openfga-test-postgres:5432/openfga?sslmode=disable" + +migration: + enabled: true + serviceAccount: + create: true + +extraObjects: + - apiVersion: v1 + kind: Secret + metadata: + name: openfga-test-postgres-creds + stringData: + POSTGRES_USER: openfga + POSTGRES_PASSWORD: changeme + POSTGRES_DB: openfga + - apiVersion: apps/v1 + kind: Deployment + metadata: + name: openfga-test-postgres + spec: + replicas: 0 # Start with Postgres DOWN + selector: + matchLabels: + app: openfga-test-postgres + template: + metadata: + labels: + app: openfga-test-postgres + spec: + containers: + - name: postgres + image: postgres:17 + ports: + - containerPort: 5432 + envFrom: + - secretRef: + name: openfga-test-postgres-creds + volumeMounts: + - name: data + mountPath: /var/lib/postgresql/data + volumes: + - name: data + emptyDir: {} + - apiVersion: v1 + kind: Service + metadata: + name: openfga-test-postgres + spec: + selector: + app: openfga-test-postgres + ports: + - port: 5432 + targetPort: 5432 diff --git a/operator/tests/values-happy-path.yaml b/operator/tests/values-happy-path.yaml new file mode 100644 index 00000000..77e6306f --- /dev/null +++ b/operator/tests/values-happy-path.yaml @@ -0,0 +1,70 @@ +# Local test values for operator-managed migration on Rancher Desktop +operator: + enabled: true + +openfga-operator: + image: + repository: openfga/openfga-operator + tag: dev + pullPolicy: Never + resources: + requests: + cpu: 10m + memory: 64Mi + +datastore: + engine: postgres + uri: "postgres://openfga:changeme@openfga-test-postgres:5432/openfga?sslmode=disable" + +migration: + enabled: true + serviceAccount: + create: true + +extraObjects: + - apiVersion: v1 + kind: Secret + metadata: + name: openfga-test-postgres-creds + stringData: + POSTGRES_USER: openfga + POSTGRES_PASSWORD: changeme + POSTGRES_DB: openfga + - apiVersion: apps/v1 + kind: Deployment + metadata: + name: openfga-test-postgres + spec: + replicas: 1 + selector: + matchLabels: + app: openfga-test-postgres + template: + metadata: + labels: + app: openfga-test-postgres + spec: + containers: + - name: postgres + image: postgres:17 + ports: + - containerPort: 5432 + envFrom: + - secretRef: + name: openfga-test-postgres-creds + volumeMounts: + - name: data + mountPath: /var/lib/postgresql/data + volumes: + - name: data + emptyDir: {} + - apiVersion: v1 + kind: Service + metadata: + name: openfga-test-postgres + spec: + selector: + app: openfga-test-postgres + ports: + - port: 5432 + targetPort: 5432 diff --git a/operator/tests/values-no-db.yaml b/operator/tests/values-no-db.yaml new file mode 100644 index 00000000..2d1cd762 --- /dev/null +++ b/operator/tests/values-no-db.yaml @@ -0,0 +1,23 @@ +# Test values with NO postgres — simulates database unavailable +operator: + enabled: true + +openfga-operator: + image: + repository: openfga/openfga-operator + tag: dev + pullPolicy: Never + resources: + requests: + cpu: 10m + memory: 64Mi + +datastore: + engine: postgres + # Points to a service that doesn't exist + uri: "postgres://openfga:changeme@postgres-does-not-exist:5432/openfga?sslmode=disable" + +migration: + enabled: true + serviceAccount: + create: true From 22ee2150deb887c575441d24449816b17124eee9 Mon Sep 17 00:00:00 2001 From: Ed Milic <edmilic@gmail.com> Date: Sat, 11 Apr 2026 08:34:55 -0400 Subject: [PATCH 03/70] fix: address PR #309 review feedback from Copilot and CodeRabbit - Harden pod security (runAsNonRoot, seccompProfile, drop ALL caps) - Find container by name instead of index to handle sidecars - Skip migration for memory datastore - Persist retry-after annotation before Job deletion to survive re-enqueue - Clear MigrationFailed condition on success - Propagate imagePullSecrets and securityContext to migration Jobs - Remove unused RBAC rules (secrets, serviceaccounts) - Add POD_NAMESPACE downward API for namespace-scoped watch default - Remove no-op migration values (timeout, backoffLimit, resources) - Fix migration SA helper to require name when create=false - Guard operator logic on both operator.enabled and migration.enabled - Build and load operator image into kind for chart-testing CI - Add path filters to operator workflow - Fix ADR inaccuracies (retry strategy, default-enabled wording) - Pin Dockerfile base image to golang:1.26.2 --- .github/workflows/operator.yml | 4 + .github/workflows/test.yml | 7 + charts/openfga-operator/templates/NOTES.txt | 7 +- .../templates/clusterrole.yaml | 6 - .../templates/deployment.yaml | 11 +- charts/openfga-operator/values.yaml | 24 ++-- charts/openfga/templates/_helpers.tpl | 6 +- charts/openfga/templates/deployment.yaml | 11 +- charts/openfga/values.schema.json | 20 --- charts/openfga/values.yaml | 9 +- docs/adr/002-operator-managed-migrations.md | 2 +- docs/adr/004-operator-deployment-model.md | 8 +- docs/adr/README.md | 2 +- operator/Dockerfile | 2 +- operator/Makefile | 1 + operator/README.md | 20 +-- operator/cmd/main.go | 13 +- operator/go.mod | 2 +- operator/internal/controller/helpers.go | 34 +++-- .../controller/migration_controller.go | 98 ++++++++++++-- .../controller/migration_controller_test.go | 128 +++++++++++++++++- 21 files changed, 312 insertions(+), 103 deletions(-) diff --git a/.github/workflows/operator.yml b/.github/workflows/operator.yml index 71f9e7f1..dc88c559 100644 --- a/.github/workflows/operator.yml +++ b/.github/workflows/operator.yml @@ -6,9 +6,13 @@ on: - main paths: - "operator/**" + - "charts/openfga-operator/**" + - ".github/workflows/operator.yml" pull_request: paths: - "operator/**" + - "charts/openfga-operator/**" + - ".github/workflows/operator.yml" workflow_dispatch: inputs: push_image: diff --git a/.github/workflows/test.yml b/.github/workflows/test.yml index 49030831..f2369c54 100644 --- a/.github/workflows/test.yml +++ b/.github/workflows/test.yml @@ -59,6 +59,13 @@ jobs: if: steps.list-changed.outputs.changed == 'true' uses: helm/kind-action@v1.14.0 + - name: Build and load operator image into kind + if: steps.list-changed.outputs.changed == 'true' + run: | + version=$(grep '^appVersion:' charts/openfga-operator/Chart.yaml | awk '{print $2}' | tr -d '"') + docker build -t "openfga/openfga-operator:${version}" operator/ + kind load docker-image "openfga/openfga-operator:${version}" --name chart-testing + - name: Run chart-testing (install) if: steps.list-changed.outputs.changed == 'true' run: ct install --target-branch ${{ github.event.repository.default_branch }} diff --git a/charts/openfga-operator/templates/NOTES.txt b/charts/openfga-operator/templates/NOTES.txt index 8c398b1c..c2e09f9a 100644 --- a/charts/openfga-operator/templates/NOTES.txt +++ b/charts/openfga-operator/templates/NOTES.txt @@ -1,11 +1,10 @@ The openfga-operator has been deployed. -NOTE: The operator container image ({{ .Values.image.repository }}:{{ .Values.image.tag | default .Chart.AppVersion }}) -does not exist yet. The operator pod will remain in ImagePullBackOff until -the Go binary is built and pushed. +NOTE: Ensure the operator image ({{ .Values.image.repository }}:{{ .Values.image.tag | default .Chart.AppVersion }}) is available in your registry. +If unavailable, the operator pod may remain in ImagePullBackOff until the image is pushed. To check operator status: kubectl get deployment --namespace {{ include "openfga-operator.namespace" . }} {{ include "openfga-operator.fullname" . }} -To view operator logs (once the image is available): +To view operator logs: kubectl logs --namespace {{ include "openfga-operator.namespace" . }} -l "app.kubernetes.io/name={{ include "openfga-operator.name" . }}" diff --git a/charts/openfga-operator/templates/clusterrole.yaml b/charts/openfga-operator/templates/clusterrole.yaml index 09d0fd7f..7dbfddf3 100644 --- a/charts/openfga-operator/templates/clusterrole.yaml +++ b/charts/openfga-operator/templates/clusterrole.yaml @@ -17,12 +17,6 @@ rules: - apiGroups: [""] resources: ["configmaps"] verbs: ["get", "list", "watch", "create", "update"] - - apiGroups: [""] - resources: ["secrets"] - verbs: ["get"] - - apiGroups: [""] - resources: ["serviceaccounts"] - verbs: ["get", "list", "create"] - apiGroups: ["coordination.k8s.io"] resources: ["leases"] verbs: ["get", "list", "watch", "create", "update"] diff --git a/charts/openfga-operator/templates/deployment.yaml b/charts/openfga-operator/templates/deployment.yaml index ae8af0d5..45706601 100644 --- a/charts/openfga-operator/templates/deployment.yaml +++ b/charts/openfga-operator/templates/deployment.yaml @@ -40,11 +40,16 @@ spec: {{- if .Values.leaderElection.enabled }} - --leader-elect {{- end }} - {{- if .Values.watchNamespace }} - - --watch-namespace={{ .Values.watchNamespace }} - {{- else if .Values.watchAllNamespaces }} + {{- if .Values.watchAllNamespaces }} - --watch-all-namespaces + {{- else if .Values.watchNamespace }} + - --watch-namespace={{ .Values.watchNamespace }} {{- end }} + env: + - name: POD_NAMESPACE + valueFrom: + fieldRef: + fieldPath: metadata.namespace ports: - name: healthz containerPort: 8081 diff --git a/charts/openfga-operator/values.yaml b/charts/openfga-operator/values.yaml index 891ad574..59c9dfba 100644 --- a/charts/openfga-operator/values.yaml +++ b/charts/openfga-operator/values.yaml @@ -21,18 +21,18 @@ serviceAccount: podAnnotations: {} -podSecurityContext: {} - # runAsNonRoot: true - # seccompProfile: - # type: RuntimeDefault - -securityContext: {} - # capabilities: - # drop: - # - ALL - # readOnlyRootFilesystem: true - # runAsNonRoot: true - # runAsUser: 65532 +podSecurityContext: + runAsNonRoot: true + seccompProfile: + type: RuntimeDefault + +securityContext: + capabilities: + drop: + - ALL + readOnlyRootFilesystem: true + runAsNonRoot: true + runAsUser: 65532 # -- Constrain the operator to watch a single namespace. # Leave empty to default to the release namespace. diff --git a/charts/openfga/templates/_helpers.tpl b/charts/openfga/templates/_helpers.tpl index 35ad94a9..cc50e03d 100644 --- a/charts/openfga/templates/_helpers.tpl +++ b/charts/openfga/templates/_helpers.tpl @@ -78,10 +78,10 @@ Create the name of the service account to use Create the name of the migration service account to use (operator mode only) */}} {{- define "openfga.migrationServiceAccountName" -}} -{{- if .Values.migration.serviceAccount.name }} -{{- .Values.migration.serviceAccount.name | trunc 63 | trimSuffix "-" }} +{{- if .Values.migration.serviceAccount.create }} +{{- default (printf "%s-migration" (include "openfga.fullname" .)) .Values.migration.serviceAccount.name | trunc 63 | trimSuffix "-" }} {{- else }} -{{- printf "%s-migration" (include "openfga.fullname" .) | trunc 63 | trimSuffix "-" }} +{{- required "migration.serviceAccount.name must be set when migration.serviceAccount.create=false" .Values.migration.serviceAccount.name | trunc 63 | trimSuffix "-" }} {{- end }} {{- end }} diff --git a/charts/openfga/templates/deployment.yaml b/charts/openfga/templates/deployment.yaml index 1cd20acf..3d101c08 100644 --- a/charts/openfga/templates/deployment.yaml +++ b/charts/openfga/templates/deployment.yaml @@ -5,15 +5,17 @@ metadata: labels: {{- include "openfga.labels" . | nindent 4 }} annotations: - {{- if .Values.operator.enabled }} + {{- if and .Values.operator.enabled .Values.migration.enabled }} openfga.dev/desired-replicas: "{{ ternary 1 .Values.replicaCount (eq .Values.datastore.engine "memory") }}" + {{- if or .Values.migration.serviceAccount.create .Values.migration.serviceAccount.name }} openfga.dev/migration-service-account: "{{ include "openfga.migrationServiceAccountName" . }}" {{- end }} + {{- end }} {{- with .Values.annotations }} {{- toYaml . | nindent 4 }} {{- end }} spec: - {{- if .Values.operator.enabled }} + {{- if and .Values.operator.enabled .Values.migration.enabled }} {{- if .Values.autoscaling.enabled }} {{- fail "operator.enabled and autoscaling.enabled cannot both be true" }} {{- end }} @@ -46,8 +48,10 @@ spec: serviceAccountName: {{ include "openfga.serviceAccountName" . }} securityContext: {{- toYaml .Values.podSecurityContext | nindent 8 }} - {{ if and (not .Values.operator.enabled) (or (and (has .Values.datastore.engine (list "postgres" "mysql")) .Values.datastore.applyMigrations .Values.datastore.waitForMigrations) .Values.extraInitContainers) }} + {{- $operatorMigration := and .Values.operator.enabled .Values.migration.enabled }} + {{ if or (and (not $operatorMigration) (or (and (has .Values.datastore.engine (list "postgres" "mysql")) .Values.datastore.applyMigrations .Values.datastore.waitForMigrations))) .Values.extraInitContainers }} initContainers: + {{- if not $operatorMigration }} {{- if and (has .Values.datastore.engine (list "postgres" "mysql")) .Values.datastore.applyMigrations .Values.datastore.waitForMigrations (eq .Values.datastore.migrationType "job") }} - name: wait-for-migration securityContext: @@ -86,6 +90,7 @@ spec: {{- include "common.tplvalues.render" ( dict "value" .Values.migrate.sidecars "context" $) | nindent 8 }} {{- end }} {{- end }} + {{- end }} {{- with .Values.extraInitContainers }} {{- toYaml . | nindent 8 }} {{- end }} diff --git a/charts/openfga/values.schema.json b/charts/openfga/values.schema.json index 76b2cf9b..54e737b7 100644 --- a/charts/openfga/values.schema.json +++ b/charts/openfga/values.schema.json @@ -1314,21 +1314,6 @@ "description": "Enable operator-managed database migrations", "default": true }, - "timeout": { - "type": ["string", "null"], - "description": "Timeout passed to the migration Job as activeDeadlineSeconds", - "default": "" - }, - "backoffLimit": { - "type": "integer", - "description": "Number of retries before marking the migration as failed", - "default": 3 - }, - "ttlSecondsAfterFinished": { - "type": "integer", - "description": "Seconds to keep completed/failed migration Jobs before cleanup", - "default": 600 - }, "serviceAccount": { "type": "object", "properties": { @@ -1351,11 +1336,6 @@ "default": "" } } - }, - "resources": { - "type": "object", - "description": "Resource requests/limits for migration Job pods", - "default": {} } } } diff --git a/charts/openfga/values.yaml b/charts/openfga/values.yaml index e843563a..24723c26 100644 --- a/charts/openfga/values.yaml +++ b/charts/openfga/values.yaml @@ -394,13 +394,8 @@ operator: # -- migration controls operator-driven migration behavior. # Only used when operator.enabled is true. migration: + # -- Enable operator-managed migrations. Set to false if you manage migrations externally. enabled: true - # -- Timeout passed to the migration Job as activeDeadlineSeconds. - timeout: "" - # -- Number of retries before marking the migration as failed. - backoffLimit: 3 - # -- Seconds to keep completed/failed migration Jobs before cleanup. - ttlSecondsAfterFinished: 600 serviceAccount: # -- Create a dedicated service account for migration Jobs. create: true @@ -409,8 +404,6 @@ migration: # -- The name of the migration service account. # If not set and create is true, defaults to {fullname}-migration. name: "" - # -- Resource requests/limits for migration Job pods. - resources: {} ## Example: Deploy a PostgreSQL instance for dev/test using official Docker images. ## For production, use a managed database service or an operator like CloudnativePG. ## Configure the chart to use the secret: diff --git a/docs/adr/002-operator-managed-migrations.md b/docs/adr/002-operator-managed-migrations.md index 8fb0cd72..9b86d9d7 100644 --- a/docs/adr/002-operator-managed-migrations.md +++ b/docs/adr/002-operator-managed-migrations.md @@ -121,7 +121,7 @@ The Job created by the operator has no Helm hook annotations. It is a standard K | Job fails | Operator sets `MigrationFailed` condition on Deployment. Does NOT scale up. User inspects Job logs. | | Job hangs | `activeDeadlineSeconds` (default 300s) kills it. Operator sees failure. | | Operator crashes | On restart, re-reads ConfigMap and Job status. Resumes from where it left off. | -| Database unreachable | Job fails to connect. Operator retries on next reconciliation (exponential backoff). | +| Database unreachable | Job fails to connect. After exhausting `backoffLimit`, operator deletes the failed Job, sets a `retry-after` annotation, and recreates a fresh Job after a fixed 60-second cooldown. Cycle repeats until the database becomes available. | ### Sequence Comparison diff --git a/docs/adr/004-operator-deployment-model.md b/docs/adr/004-operator-deployment-model.md index bedb12c1..745e7775 100644 --- a/docs/adr/004-operator-deployment-model.md +++ b/docs/adr/004-operator-deployment-model.md @@ -35,7 +35,7 @@ The operator Deployment, RBAC, and CRDs are templates in the main OpenFGA chart. **C. Operator as a conditional subchart dependency (selected)** -The operator is a separate Helm chart (`openfga-operator`) that the main chart declares as a conditional dependency. Enabled by default, but users can disable it. +The operator is a separate Helm chart (`openfga-operator`) that the main chart declares as a conditional dependency. Disabled by default for backward compatibility; users opt in with `operator.enabled: true`. *Example:* ```bash @@ -71,7 +71,7 @@ helm-charts/ ├── charts/ │ ├── openfga/ # Main chart (existing) │ │ ├── Chart.yaml # Declares openfga-operator as dependency -│ │ ├── values.yaml # operator.enabled: true +│ │ ├── values.yaml # operator.enabled: false (opt-in) │ │ ├── templates/ │ │ └── crds/ # Empty in Stage 1 │ │ @@ -124,8 +124,8 @@ kubectl apply -f https://github.com/openfga/helm-charts/releases/download/v0.2.0 | Mode | Command | Use case | |------|---------|----------| -| **All-in-one** (default) | `helm install openfga openfga/openfga` | Most users. Single install, operator included. | -| **Operator disabled** | `helm install openfga openfga/openfga --set operator.enabled=false` | Operator managed separately or not needed (memory datastore). | +| **Default** (no operator) | `helm install openfga openfga/openfga` | Backward compatible. Uses Helm hooks for migration. | +| **All-in-one** | `helm install openfga openfga/openfga --set operator.enabled=true` | Single install with operator-managed migrations. | | **Operator standalone** | `helm install op openfga/openfga-operator -n openfga-system` | Cluster-wide operator serving multiple OpenFGA instances. | ### Multi-Instance Considerations diff --git a/docs/adr/README.md b/docs/adr/README.md index 298a9e32..5f805122 100644 --- a/docs/adr/README.md +++ b/docs/adr/README.md @@ -30,7 +30,7 @@ ADRs are **immutable once accepted** — if a decision changes, you write a new ## ADR Lifecycle -``` +```text Proposed → Accepted → (optionally) Superseded or Deprecated ↑ │ feedback loop diff --git a/operator/Dockerfile b/operator/Dockerfile index be9097eb..7d836a3a 100644 --- a/operator/Dockerfile +++ b/operator/Dockerfile @@ -1,4 +1,4 @@ -FROM golang:1.25 AS builder +FROM golang:1.26.2 AS builder WORKDIR /workspace COPY go.mod go.sum ./ diff --git a/operator/Makefile b/operator/Makefile index 575bb7cd..4b97c0cb 100644 --- a/operator/Makefile +++ b/operator/Makefile @@ -3,6 +3,7 @@ IMG ?= openfga/openfga-operator:dev .PHONY: build test vet fmt lint docker-build docker-push clean build: + mkdir -p bin go build -o bin/operator ./cmd/ test: diff --git a/operator/README.md b/operator/README.md index a2b06996..8e576db8 100644 --- a/operator/README.md +++ b/operator/README.md @@ -2,12 +2,12 @@ A Kubernetes operator that manages database migrations for OpenFGA deployments. Instead of relying on Helm hooks and init containers, the operator watches OpenFGA Deployments, detects version changes, and orchestrates migrations as regular Jobs. -This is **Stage 1** of the operator — focused solely on migration orchestration. See [ADR-001](../docs/adr/001-adopt-operator.md) for the full roadmap. +This is **Stage 1** of the operator — focused solely on migration orchestration. See [ADR-001](../docs/adr/001-adopt-openfga-operator.md) for the full roadmap. ## How It Works 1. The operator watches Deployments labeled `app.kubernetes.io/part-of: openfga` -2. When a version change is detected (comparing the container image tag to the `openfga-migration-status` ConfigMap), the operator: +2. When a version change is detected (comparing the container image tag to the `{name}-migration-status` ConfigMap), the operator: - Keeps the Deployment at 0 replicas - Creates a migration Job running `openfga migrate` - Waits for the Job to complete @@ -108,14 +108,14 @@ The operator accepts the following flags: | Flag | Default | Description | |------|---------|-------------| -| `--leader-elect` | `false` | Enable leader election | -| `--watch-namespace` | `""` | Namespace to watch (defaults to release namespace) | -| `--watch-all-namespaces` | `false` | Watch all namespaces | -| `--metrics-bind-address` | `:8080` | Metrics endpoint address | -| `--health-probe-bind-address` | `:8081` | Health probe endpoint address | -| `--backoff-limit` | `3` | BackoffLimit for migration Jobs | -| `--active-deadline-seconds` | `300` | ActiveDeadlineSeconds for migration Jobs | -| `--ttl-seconds-after-finished` | `300` | TTLSecondsAfterFinished for migration Jobs | +| `--leader-elect` | `false` | Enable leader election so only one replica actively reconciles at a time. Required when running multiple operator replicas for high availability; standby pods wait for the leader's Lease to expire before taking over. Not needed for single-replica deployments. | +| `--watch-namespace` | `""` | Namespace to watch for OpenFGA Deployments. Defaults to the operator pod's own namespace (via `POD_NAMESPACE` env var). Set explicitly for multi-namespace setups. | +| `--watch-all-namespaces` | `false` | Watch all namespaces for OpenFGA Deployments, making the operator cluster-wide. Overrides `--watch-namespace`. | +| `--metrics-bind-address` | `:8080` | Address the Prometheus metrics endpoint binds to. Change only if the default port conflicts with other containers in the pod. | +| `--health-probe-bind-address` | `:8081` | Address the Kubernetes liveness and readiness probe endpoints bind to. Change only if the default port conflicts. | +| `--backoff-limit` | `3` | Number of times a migration Job's pod can fail before the Job is considered failed. After hitting this limit the operator deletes the Job, sets a `MigrationFailed` condition on the Deployment, and retries after a 60-second cooldown. | +| `--active-deadline-seconds` | `300` | Maximum wall-clock seconds a migration Job can run before Kubernetes terminates it. Prevents stuck migrations from blocking the pipeline indefinitely. Increase for very large databases. | +| `--ttl-seconds-after-finished` | `300` | Seconds Kubernetes keeps a completed or failed Job (and its pods) before garbage-collecting them, giving you time to inspect logs. | When deployed via the Helm subchart, these are configured through `values.yaml`. See `charts/openfga-operator/values.yaml` for all available options. diff --git a/operator/cmd/main.go b/operator/cmd/main.go index e4c9d3ea..9fb91070 100644 --- a/operator/cmd/main.go +++ b/operator/cmd/main.go @@ -5,7 +5,6 @@ import ( "os" "k8s.io/apimachinery/pkg/runtime" - utilruntime "k8s.io/apimachinery/pkg/runtime/serializer" clientgoscheme "k8s.io/client-go/kubernetes/scheme" ctrl "sigs.k8s.io/controller-runtime" "sigs.k8s.io/controller-runtime/pkg/cache" @@ -20,8 +19,6 @@ var scheme = runtime.NewScheme() func init() { _ = clientgoscheme.AddToScheme(scheme) - // Suppress unused import. - _ = utilruntime.CodecFactory{} } func main() { @@ -37,7 +34,7 @@ func main() { ) flag.BoolVar(&leaderElect, "leader-elect", false, "Enable leader election for the controller manager.") - flag.StringVar(&watchNamespace, "watch-namespace", "", "Namespace to watch. Defaults to the release namespace.") + flag.StringVar(&watchNamespace, "watch-namespace", "", "Namespace to watch. Defaults to the operator pod namespace.") flag.BoolVar(&watchAllNamespaces, "watch-all-namespaces", false, "Watch all namespaces.") flag.StringVar(&metricsAddr, "metrics-bind-address", ":8080", "The address the metric endpoint binds to.") flag.StringVar(&healthProbeAddr, "health-probe-bind-address", ":8081", "The address the health probe endpoint binds to.") @@ -52,6 +49,14 @@ func main() { ctrl.SetLogger(zap.New(zap.UseFlagOptions(&opts))) logger := ctrl.Log.WithName("setup") + // Fall back to the pod's namespace when no explicit scope is set. + if !watchAllNamespaces && watchNamespace == "" { + if podNS, ok := os.LookupEnv("POD_NAMESPACE"); ok && podNS != "" { + watchNamespace = podNS + logger.Info("defaulting watch scope to pod namespace", "namespace", podNS) + } + } + // Configure cache namespace restrictions. var cacheOpts cache.Options if watchNamespace != "" && !watchAllNamespaces { diff --git a/operator/go.mod b/operator/go.mod index fea96493..8cd2bbe4 100644 --- a/operator/go.mod +++ b/operator/go.mod @@ -1,6 +1,6 @@ module github.com/openfga/openfga-operator -go 1.25.6 +go 1.26.2 require ( k8s.io/api v0.35.3 diff --git a/operator/internal/controller/helpers.go b/operator/internal/controller/helpers.go index 175805d0..23e0d8dc 100644 --- a/operator/internal/controller/helpers.go +++ b/operator/internal/controller/helpers.go @@ -26,6 +26,7 @@ const ( // Annotations set on the Deployment by the Helm chart / operator. AnnotationDesiredReplicas = "openfga.dev/desired-replicas" AnnotationMigrationServiceAccount = "openfga.dev/migration-service-account" + AnnotationRetryAfter = "openfga.dev/migration-retry-after" // Defaults for migration Job configuration. DefaultBackoffLimit int32 = 3 @@ -68,17 +69,29 @@ func migrationJobName(deploymentName string) string { return deploymentName + "-migrate" } -// buildMigrationJob constructs a migration Job for the given Deployment and version. +// findOpenFGAContainer finds the OpenFGA container in the Deployment's pod spec. +// It looks for a container named "openfga" first, then falls back to the first container. +func findOpenFGAContainer(deployment *appsv1.Deployment) *corev1.Container { + for i := range deployment.Spec.Template.Spec.Containers { + if deployment.Spec.Template.Spec.Containers[i].Name == "openfga" { + return &deployment.Spec.Template.Spec.Containers[i] + } + } + // Fallback: use the first container (for charts that don't name it "openfga"). + if len(deployment.Spec.Template.Spec.Containers) > 0 { + return &deployment.Spec.Template.Spec.Containers[0] + } + return nil +} + +// buildMigrationJob constructs a migration Job for the given Deployment. func buildMigrationJob( deployment *appsv1.Deployment, - desiredVersion string, + mainContainer *corev1.Container, backoffLimit int32, activeDeadlineSeconds int64, ttlSecondsAfterFinished int32, ) *batchv1.Job { - // Extract the main container's image and datastore env vars. - mainContainer := deployment.Spec.Template.Spec.Containers[0] - // Determine the migration service account. migrationSA := deployment.Annotations[AnnotationMigrationServiceAccount] if migrationSA == "" { @@ -127,12 +140,15 @@ func buildMigrationJob( Spec: corev1.PodSpec{ ServiceAccountName: migrationSA, RestartPolicy: corev1.RestartPolicyNever, + ImagePullSecrets: deployment.Spec.Template.Spec.ImagePullSecrets, + SecurityContext: deployment.Spec.Template.Spec.SecurityContext, Containers: []corev1.Container{ { - Name: "migrate-database", - Image: mainContainer.Image, - Args: []string{"migrate"}, - Env: datastoreEnvVars, + Name: "migrate-database", + Image: mainContainer.Image, + Args: []string{"migrate"}, + Env: datastoreEnvVars, + SecurityContext: mainContainer.SecurityContext, }, }, // Inherit scheduling constraints from the parent Deployment. diff --git a/operator/internal/controller/migration_controller.go b/operator/internal/controller/migration_controller.go index 935e4ffc..9da395cc 100644 --- a/operator/internal/controller/migration_controller.go +++ b/operator/internal/controller/migration_controller.go @@ -3,6 +3,7 @@ package controller import ( "context" "fmt" + "strings" "time" appsv1 "k8s.io/api/apps/v1" @@ -46,14 +47,24 @@ func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) ( return ctrl.Result{}, err } - // 2. Extract the desired version from the Deployment's image tag. - if len(deployment.Spec.Template.Spec.Containers) == 0 { + // 2. Find the OpenFGA container and extract the desired version. + mainContainer := findOpenFGAContainer(deployment) + if mainContainer == nil { logger.Info("deployment has no containers, skipping") return ctrl.Result{}, nil } - desiredVersion := extractImageTag(deployment.Spec.Template.Spec.Containers[0].Image) + desiredVersion := extractImageTag(mainContainer.Image) - // 3. Check current migration status from ConfigMap. + // 3. Skip migration for memory datastore — just ensure the Deployment is scaled up. + if isMemoryDatastore(mainContainer) { + logger.V(1).Info("memory datastore detected, skipping migration") + if _, scaleErr := ensureDeploymentScaled(ctx, r.Client, deployment); scaleErr != nil { + return ctrl.Result{}, scaleErr + } + return ctrl.Result{}, nil + } + + // 4. Check current migration status from ConfigMap. configMap := &corev1.ConfigMap{} cmName := migrationConfigMapName(req.Name) err := r.Get(ctx, types.NamespacedName{Name: cmName, Namespace: req.Namespace}, configMap) @@ -65,9 +76,13 @@ func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) ( return ctrl.Result{}, fmt.Errorf("getting migration status: %w", err) } - // 4. If versions match, ensure Deployment is scaled up and return. + // 5. If versions match, ensure Deployment is scaled up and return. if currentVersion == desiredVersion { logger.V(1).Info("migration up to date", "version", desiredVersion) + clearMigrationFailedCondition(deployment) + if patchErr := r.Status().Update(ctx, deployment); patchErr != nil { + logger.Error(patchErr, "failed to clear MigrationFailed condition") + } if _, scaleErr := ensureDeploymentScaled(ctx, r.Client, deployment); scaleErr != nil { return ctrl.Result{}, scaleErr } @@ -76,12 +91,22 @@ func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) ( logger.Info("migration needed", "currentVersion", currentVersion, "desiredVersion", desiredVersion) - // 5. Ensure the Deployment is scaled to zero before migrating. + // 6. Ensure the Deployment is scaled to zero before migrating. if err := scaleDeploymentToZero(ctx, r.Client, deployment); err != nil { return ctrl.Result{}, err } - // 6. Check if a migration Job already exists. + // 7. Check retry-after annotation to honor backoff cooldown. + if retryAfter, ok := deployment.Annotations[AnnotationRetryAfter]; ok { + retryTime, parseErr := time.Parse(time.RFC3339, retryAfter) + if parseErr == nil && time.Now().Before(retryTime) { + remaining := time.Until(retryTime) + logger.V(1).Info("in retry cooldown", "retryAfter", retryAfter, "remaining", remaining) + return ctrl.Result{RequeueAfter: remaining}, nil + } + } + + // 8. Check if a migration Job already exists. jobName := migrationJobName(req.Name) job := &batchv1.Job{} err = r.Get(ctx, types.NamespacedName{Name: jobName, Namespace: req.Namespace}, job) @@ -90,11 +115,19 @@ func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) ( // Create the migration Job. job = buildMigrationJob( deployment, - desiredVersion, + mainContainer, r.BackoffLimit, r.ActiveDeadlineSeconds, r.TTLSecondsAfterFinished, ) + // Clear the retry-after annotation now that we're creating a new Job. + if _, hasRetry := deployment.Annotations[AnnotationRetryAfter]; hasRetry { + patch := client.MergeFrom(deployment.DeepCopy()) + delete(deployment.Annotations, AnnotationRetryAfter) + if patchErr := r.Patch(ctx, deployment, patch); patchErr != nil { + logger.Error(patchErr, "failed to clear retry-after annotation") + } + } if createErr := r.Create(ctx, job); createErr != nil { return ctrl.Result{}, fmt.Errorf("creating migration job: %w", createErr) } @@ -104,10 +137,16 @@ func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) ( return ctrl.Result{}, fmt.Errorf("getting migration job: %w", err) } - // 7. Check Job status. + // 9. Check Job status. if job.Status.Succeeded >= 1 { logger.Info("migration succeeded", "version", desiredVersion) + // Clear MigrationFailed condition. + clearMigrationFailedCondition(deployment) + if patchErr := r.Status().Update(ctx, deployment); patchErr != nil { + logger.Error(patchErr, "failed to clear MigrationFailed condition") + } + // Update migration status ConfigMap. if statusErr := updateMigrationStatus(ctx, r.Client, deployment, desiredVersion, jobName); statusErr != nil { return ctrl.Result{}, statusErr @@ -135,8 +174,19 @@ func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) ( logger.Error(patchErr, "failed to set MigrationFailed condition") } + // Persist a retry-after annotation so the cooldown is honored even + // when the Job deletion triggers an immediate re-enqueue. + retryAfter := time.Now().Add(60 * time.Second).UTC().Format(time.RFC3339) + patch := client.MergeFrom(deployment.DeepCopy()) + if deployment.Annotations == nil { + deployment.Annotations = make(map[string]string) + } + deployment.Annotations[AnnotationRetryAfter] = retryAfter + if patchErr := r.Patch(ctx, deployment, patch); patchErr != nil { + logger.Error(patchErr, "failed to set retry-after annotation") + } + // Delete the failed Job so a fresh one is created on the next reconcile. - // This allows auto-recovery when the database comes back. propagation := metav1.DeletePropagationBackground if delErr := r.Delete(ctx, job, &client.DeleteOptions{ PropagationPolicy: &propagation, @@ -145,15 +195,26 @@ func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) ( } logger.Info("deleted failed migration job, will retry", "job", jobName) - // Requeue after a longer delay to avoid tight retry loops. + // Requeue after the cooldown period. return ctrl.Result{RequeueAfter: 60 * time.Second}, nil } - // 8. Job still running — requeue. + // 10. Job still running — requeue. logger.V(1).Info("migration job in progress", "job", jobName) return ctrl.Result{RequeueAfter: 10 * time.Second}, nil } +// isMemoryDatastore checks if the Deployment is using the memory datastore +// (no database migration needed). +func isMemoryDatastore(container *corev1.Container) bool { + for _, env := range container.Env { + if env.Name == "OPENFGA_DATASTORE_ENGINE" { + return strings.EqualFold(env.Value, "memory") + } + } + return false +} + // setMigrationFailedCondition sets a MigrationFailed condition on the Deployment. func setMigrationFailedCondition(deployment *appsv1.Deployment, version string) { condition := appsv1.DeploymentCondition{ @@ -174,6 +235,19 @@ func setMigrationFailedCondition(deployment *appsv1.Deployment, version string) deployment.Status.Conditions = append(deployment.Status.Conditions, condition) } +// clearMigrationFailedCondition removes or sets the MigrationFailed condition to False. +func clearMigrationFailedCondition(deployment *appsv1.Deployment) { + for i, c := range deployment.Status.Conditions { + if c.Type == "MigrationFailed" { + deployment.Status.Conditions[i].Status = corev1.ConditionFalse + deployment.Status.Conditions[i].LastTransitionTime = metav1.Now() + deployment.Status.Conditions[i].Reason = "MigrationSucceeded" + deployment.Status.Conditions[i].Message = "Migration completed successfully." + return + } + } +} + // SetupWithManager sets up the controller with the Manager. func (r *MigrationReconciler) SetupWithManager(mgr ctrl.Manager) error { // Only watch Deployments that are part of OpenFGA. diff --git a/operator/internal/controller/migration_controller_test.go b/operator/internal/controller/migration_controller_test.go index dc0f0c56..f463c7d9 100644 --- a/operator/internal/controller/migration_controller_test.go +++ b/operator/internal/controller/migration_controller_test.go @@ -230,7 +230,7 @@ func TestReconcile_JobSucceeded_UpdatesConfigMapAndScalesUp(t *testing.T) { } } -func TestReconcile_JobFailed_DeletesJobAndRequeues(t *testing.T) { +func TestReconcile_JobFailed_SetsRetryAnnotationAndRequeues(t *testing.T) { // Given: a Deployment at 0 replicas and a failed migration Job. dep := newTestDeployment("openfga", "default", "openfga/openfga:v1.14.0", 0) dep.Annotations[AnnotationDesiredReplicas] = "3" @@ -295,6 +295,87 @@ func TestReconcile_JobFailed_DeletesJobAndRequeues(t *testing.T) { }, deletedJob); getErr == nil { t.Error("expected failed migration job to be deleted") } + + // Verify retry-after annotation was set on the Deployment. + if _, ok := updated.Annotations[AnnotationRetryAfter]; !ok { + t.Error("expected retry-after annotation to be set on Deployment") + } +} + +func TestReconcile_RetryAfterCooldown_SkipsJobCreation(t *testing.T) { + // Given: a Deployment with a retry-after annotation in the future. + dep := newTestDeployment("openfga", "default", "openfga/openfga:v1.14.0", 0) + dep.Annotations[AnnotationDesiredReplicas] = "3" + dep.Annotations[AnnotationRetryAfter] = time.Now().Add(30 * time.Second).UTC().Format(time.RFC3339) + + r := newReconciler(dep) + + // When: reconciling. + result, err := r.Reconcile(context.Background(), ctrl.Request{ + NamespacedName: types.NamespacedName{Name: "openfga", Namespace: "default"}, + }) + + // Then: no error, requeue with remaining cooldown time. + if err != nil { + t.Fatalf("unexpected error: %v", err) + } + if result.RequeueAfter == 0 { + t.Error("expected requeue during cooldown") + } + if result.RequeueAfter > 30*time.Second { + t.Errorf("expected requeue within 30s, got %v", result.RequeueAfter) + } + + // Verify no Job was created. + job := &batchv1.Job{} + if getErr := r.Get(context.Background(), types.NamespacedName{ + Name: "openfga-migrate", Namespace: "default", + }, job); getErr == nil { + t.Error("expected no migration job during cooldown") + } +} + +func TestReconcile_MemoryDatastore_SkipsMigration(t *testing.T) { + // Given: a Deployment using the memory datastore. + dep := newTestDeployment("openfga", "default", "openfga/openfga:v1.14.0", 0) + dep.Annotations[AnnotationDesiredReplicas] = "1" + dep.Spec.Template.Spec.Containers[0].Env = []corev1.EnvVar{ + {Name: "OPENFGA_DATASTORE_ENGINE", Value: "memory"}, + } + + r := newReconciler(dep) + + // When: reconciling. + result, err := r.Reconcile(context.Background(), ctrl.Request{ + NamespacedName: types.NamespacedName{Name: "openfga", Namespace: "default"}, + }) + + // Then: no error, no requeue. + if err != nil { + t.Fatalf("unexpected error: %v", err) + } + if result.RequeueAfter != 0 { + t.Error("expected no requeue for memory datastore") + } + + // Verify Deployment was scaled up (no migration needed). + updated := &appsv1.Deployment{} + if getErr := r.Get(context.Background(), types.NamespacedName{ + Name: "openfga", Namespace: "default", + }, updated); getErr != nil { + t.Fatalf("getting deployment: %v", getErr) + } + if *updated.Spec.Replicas != 1 { + t.Errorf("expected 1 replica, got %d", *updated.Spec.Replicas) + } + + // Verify no Job was created. + job := &batchv1.Job{} + if getErr := r.Get(context.Background(), types.NamespacedName{ + Name: "openfga-migrate", Namespace: "default", + }, job); getErr == nil { + t.Error("expected no migration job for memory datastore") + } } func TestReconcile_DeploymentNotFound_NoError(t *testing.T) { @@ -312,6 +393,51 @@ func TestReconcile_DeploymentNotFound_NoError(t *testing.T) { } } +func TestReconcile_FindContainerByName(t *testing.T) { + // Given: a Deployment with a sidecar before the openfga container. + dep := newTestDeployment("openfga", "default", "openfga/openfga:v1.14.0", 0) + dep.Spec.Template.Spec.Containers = []corev1.Container{ + { + Name: "sidecar", + Image: "envoyproxy/envoy:v1.30", + }, + { + Name: "openfga", + Image: "openfga/openfga:v1.14.0", + Env: []corev1.EnvVar{ + {Name: "OPENFGA_DATASTORE_ENGINE", Value: "postgres"}, + {Name: "OPENFGA_DATASTORE_URI", Value: "postgres://localhost/openfga"}, + }, + }, + } + + r := newReconciler(dep) + + // When: reconciling. + result, err := r.Reconcile(context.Background(), ctrl.Request{ + NamespacedName: types.NamespacedName{Name: "openfga", Namespace: "default"}, + }) + + // Then: Job should use the openfga container's image, not the sidecar's. + if err != nil { + t.Fatalf("unexpected error: %v", err) + } + if result.RequeueAfter == 0 { + t.Error("expected requeue, got none") + } + + job := &batchv1.Job{} + if err := r.Get(context.Background(), types.NamespacedName{ + Name: "openfga-migrate", Namespace: "default", + }, job); err != nil { + t.Fatalf("expected migration job to be created: %v", err) + } + + if job.Spec.Template.Spec.Containers[0].Image != "openfga/openfga:v1.14.0" { + t.Errorf("expected job image openfga/openfga:v1.14.0, got %s", job.Spec.Template.Spec.Containers[0].Image) + } +} + func TestExtractImageTag(t *testing.T) { tests := []struct { image string From 539bb14a665a1b30daf8489123198e87ae4dee86 Mon Sep 17 00:00:00 2001 From: Ed Milic <edmilic@gmail.com> Date: Sat, 11 Apr 2026 09:14:14 -0400 Subject: [PATCH 04/70] fix: address additional Copilot review feedback on PR #309 - Render extraInitContainers in operator mode (previously skipped) - Add version label to migration Jobs and delete stale Jobs on image change - Use namespaced Role/RoleBinding when watchAllNamespaces is false - Replace Status().Update with Status().Patch to avoid write conflicts - Fix logger.Error(nil, ...) to logger.Info for expected failure state - Wire desiredVersion param into buildMigrationJob for version tracking - Update RBAC: deployments/status verb from update to patch - Don't force replicas: 0 for memory engine in operator mode - Guard migration SA creation on migration.enabled - Document both required labels and mutable tag limitation in README --- .../templates/clusterrole.yaml | 9 ++++++- .../templates/clusterrolebinding.yaml | 11 +++++++++ charts/openfga/templates/deployment.yaml | 11 +++++---- charts/openfga/templates/serviceaccount.yaml | 2 +- operator/README.md | 6 ++++- operator/internal/controller/helpers.go | 2 ++ .../controller/migration_controller.go | 24 +++++++++++++++---- 7 files changed, 53 insertions(+), 12 deletions(-) diff --git a/charts/openfga-operator/templates/clusterrole.yaml b/charts/openfga-operator/templates/clusterrole.yaml index 7dbfddf3..652b48e3 100644 --- a/charts/openfga-operator/templates/clusterrole.yaml +++ b/charts/openfga-operator/templates/clusterrole.yaml @@ -1,7 +1,14 @@ apiVersion: rbac.authorization.k8s.io/v1 +{{- if .Values.watchAllNamespaces }} kind: ClusterRole +{{- else }} +kind: Role +{{- end }} metadata: name: {{ include "openfga-operator.fullname" . }} + {{- if not .Values.watchAllNamespaces }} + namespace: {{ include "openfga-operator.namespace" . }} + {{- end }} labels: {{- include "openfga-operator.labels" . | nindent 4 }} rules: @@ -10,7 +17,7 @@ rules: verbs: ["get", "list", "watch", "patch"] - apiGroups: ["apps"] resources: ["deployments/status"] - verbs: ["update"] + verbs: ["patch"] - apiGroups: ["batch"] resources: ["jobs"] verbs: ["get", "list", "watch", "create", "delete"] diff --git a/charts/openfga-operator/templates/clusterrolebinding.yaml b/charts/openfga-operator/templates/clusterrolebinding.yaml index 854521ab..cfca8d1a 100644 --- a/charts/openfga-operator/templates/clusterrolebinding.yaml +++ b/charts/openfga-operator/templates/clusterrolebinding.yaml @@ -1,12 +1,23 @@ apiVersion: rbac.authorization.k8s.io/v1 +{{- if .Values.watchAllNamespaces }} kind: ClusterRoleBinding +{{- else }} +kind: RoleBinding +{{- end }} metadata: name: {{ include "openfga-operator.fullname" . }} + {{- if not .Values.watchAllNamespaces }} + namespace: {{ include "openfga-operator.namespace" . }} + {{- end }} labels: {{- include "openfga-operator.labels" . | nindent 4 }} roleRef: apiGroup: rbac.authorization.k8s.io + {{- if .Values.watchAllNamespaces }} kind: ClusterRole + {{- else }} + kind: Role + {{- end }} name: {{ include "openfga-operator.fullname" . }} subjects: - kind: ServiceAccount diff --git a/charts/openfga/templates/deployment.yaml b/charts/openfga/templates/deployment.yaml index 3d101c08..2de84cdb 100644 --- a/charts/openfga/templates/deployment.yaml +++ b/charts/openfga/templates/deployment.yaml @@ -19,7 +19,7 @@ spec: {{- if .Values.autoscaling.enabled }} {{- fail "operator.enabled and autoscaling.enabled cannot both be true" }} {{- end }} - replicas: 0 + replicas: {{ ternary 1 0 (eq .Values.datastore.engine "memory") }} {{- else if not .Values.autoscaling.enabled }} replicas: {{ ternary 1 .Values.replicaCount (eq .Values.datastore.engine "memory")}} {{- end }} @@ -49,10 +49,11 @@ spec: securityContext: {{- toYaml .Values.podSecurityContext | nindent 8 }} {{- $operatorMigration := and .Values.operator.enabled .Values.migration.enabled }} - {{ if or (and (not $operatorMigration) (or (and (has .Values.datastore.engine (list "postgres" "mysql")) .Values.datastore.applyMigrations .Values.datastore.waitForMigrations))) .Values.extraInitContainers }} + {{- $needsMigrationInit := and (not $operatorMigration) (has .Values.datastore.engine (list "postgres" "mysql")) .Values.datastore.applyMigrations .Values.datastore.waitForMigrations }} + {{- if or $needsMigrationInit .Values.extraInitContainers }} initContainers: - {{- if not $operatorMigration }} - {{- if and (has .Values.datastore.engine (list "postgres" "mysql")) .Values.datastore.applyMigrations .Values.datastore.waitForMigrations (eq .Values.datastore.migrationType "job") }} + {{- if $needsMigrationInit }} + {{- if eq .Values.datastore.migrationType "job" }} - name: wait-for-migration securityContext: {{- toYaml .Values.securityContext | nindent 12 }} @@ -62,7 +63,7 @@ spec: resources: {{- toYaml .Values.datastore.migrations.resources | nindent 12 }} {{- end }} - {{- if and (has .Values.datastore.engine (list "postgres" "mysql")) (eq .Values.datastore.migrationType "initContainer") }} + {{- if eq .Values.datastore.migrationType "initContainer" }} {{- with .Values.migrate.extraInitContainers }} {{- toYaml . | nindent 8 }} {{- end }} diff --git a/charts/openfga/templates/serviceaccount.yaml b/charts/openfga/templates/serviceaccount.yaml index 278be07b..f732c46e 100644 --- a/charts/openfga/templates/serviceaccount.yaml +++ b/charts/openfga/templates/serviceaccount.yaml @@ -10,7 +10,7 @@ metadata: {{- toYaml . | nindent 4 }} {{- end }} {{- end }} -{{- if and .Values.operator.enabled .Values.migration.serviceAccount.create }} +{{- if and .Values.operator.enabled .Values.migration.enabled .Values.migration.serviceAccount.create }} --- apiVersion: v1 kind: ServiceAccount diff --git a/operator/README.md b/operator/README.md index 8e576db8..a2e927b1 100644 --- a/operator/README.md +++ b/operator/README.md @@ -6,7 +6,7 @@ This is **Stage 1** of the operator — focused solely on migration orchestratio ## How It Works -1. The operator watches Deployments labeled `app.kubernetes.io/part-of: openfga` +1. The operator watches Deployments labeled `app.kubernetes.io/part-of: openfga` and `app.kubernetes.io/component: authorization-controller` 2. When a version change is detected (comparing the container image tag to the `{name}-migration-status` ConfigMap), the operator: - Keeps the Deployment at 0 replicas - Creates a migration Job running `openfga migrate` @@ -127,3 +127,7 @@ The operator reads these annotations from the OpenFGA Deployment: |------------|-------------| | `openfga.dev/desired-replicas` | The replica count to restore after migration succeeds. Set by the Helm chart. | | `openfga.dev/migration-service-account` | The ServiceAccount to use for migration Jobs. Defaults to the Deployment's SA. | + +## Limitations + +- **Mutable image tags:** The operator detects version changes by comparing the container image tag (or digest). If you deploy with a mutable tag like `latest` or reuse the same tag for different builds, the operator will not detect changes and will skip the migration. Use immutable tags (e.g., `v1.14.0`) or pin images by digest for reliable migration triggering. diff --git a/operator/internal/controller/helpers.go b/operator/internal/controller/helpers.go index 23e0d8dc..a9c37d51 100644 --- a/operator/internal/controller/helpers.go +++ b/operator/internal/controller/helpers.go @@ -88,6 +88,7 @@ func findOpenFGAContainer(deployment *appsv1.Deployment) *corev1.Container { func buildMigrationJob( deployment *appsv1.Deployment, mainContainer *corev1.Container, + desiredVersion string, backoffLimit int32, activeDeadlineSeconds int64, ttlSecondsAfterFinished int32, @@ -114,6 +115,7 @@ func buildMigrationJob( LabelPartOf: LabelPartOfValue, LabelComponent: "migration", "app.kubernetes.io/managed-by": "openfga-operator", + "app.kubernetes.io/version": desiredVersion, }, OwnerReferences: []metav1.OwnerReference{ { diff --git a/operator/internal/controller/migration_controller.go b/operator/internal/controller/migration_controller.go index 9da395cc..77dd22eb 100644 --- a/operator/internal/controller/migration_controller.go +++ b/operator/internal/controller/migration_controller.go @@ -79,8 +79,9 @@ func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) ( // 5. If versions match, ensure Deployment is scaled up and return. if currentVersion == desiredVersion { logger.V(1).Info("migration up to date", "version", desiredVersion) + statusPatch := client.MergeFrom(deployment.DeepCopy()) clearMigrationFailedCondition(deployment) - if patchErr := r.Status().Update(ctx, deployment); patchErr != nil { + if patchErr := r.Status().Patch(ctx, deployment, statusPatch); patchErr != nil { logger.Error(patchErr, "failed to clear MigrationFailed condition") } if _, scaleErr := ensureDeploymentScaled(ctx, r.Client, deployment); scaleErr != nil { @@ -116,6 +117,7 @@ func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) ( job = buildMigrationJob( deployment, mainContainer, + desiredVersion, r.BackoffLimit, r.ActiveDeadlineSeconds, r.TTLSecondsAfterFinished, @@ -137,13 +139,26 @@ func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) ( return ctrl.Result{}, fmt.Errorf("getting migration job: %w", err) } + // 8b. If the existing Job is for a different version, delete it and recreate. + if jobVersion := job.Labels["app.kubernetes.io/version"]; jobVersion != "" && jobVersion != desiredVersion { + logger.Info("existing migration job is for a different version, deleting", "jobVersion", jobVersion, "desiredVersion", desiredVersion) + propagation := metav1.DeletePropagationBackground + if delErr := r.Delete(ctx, job, &client.DeleteOptions{ + PropagationPolicy: &propagation, + }); delErr != nil && !apierrors.IsNotFound(delErr) { + return ctrl.Result{}, fmt.Errorf("deleting stale migration job: %w", delErr) + } + return ctrl.Result{RequeueAfter: 5 * time.Second}, nil + } + // 9. Check Job status. if job.Status.Succeeded >= 1 { logger.Info("migration succeeded", "version", desiredVersion) // Clear MigrationFailed condition. + statusPatch := client.MergeFrom(deployment.DeepCopy()) clearMigrationFailedCondition(deployment) - if patchErr := r.Status().Update(ctx, deployment); patchErr != nil { + if patchErr := r.Status().Patch(ctx, deployment, statusPatch); patchErr != nil { logger.Error(patchErr, "failed to clear MigrationFailed condition") } @@ -166,11 +181,12 @@ func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) ( } if job.Status.Failed >= backoffLimit { - logger.Error(nil, "migration job failed, will delete and retry", "job", jobName, "version", desiredVersion) + logger.Info("migration job failed, will delete and retry", "job", jobName, "version", desiredVersion) // Set condition so kubectl describe shows the failure. + statusPatch := client.MergeFrom(deployment.DeepCopy()) setMigrationFailedCondition(deployment, desiredVersion) - if patchErr := r.Status().Update(ctx, deployment); patchErr != nil { + if patchErr := r.Status().Patch(ctx, deployment, statusPatch); patchErr != nil { logger.Error(patchErr, "failed to set MigrationFailed condition") } From 76c5b1291996daff5796d78106ffa26c6b104e0b Mon Sep 17 00:00:00 2001 From: Ed Milic <edmilic@gmail.com> Date: Sat, 11 Apr 2026 09:57:11 -0400 Subject: [PATCH 05/70] fix: address remaining PR #309 review feedback - Add opt-in annotation (openfga.dev/migration-enabled) so the operator only manages migrations for explicitly opted-in Deployments - Propagate volumes, volumeMounts, and envFrom from the Deployment to migration Jobs for TLS certs and file-based credentials - Remove watchAllNamespaces option; operator is now always namespace-scoped - Update ADR-004 dependency example to match actual file:// reference - Add test for migration-not-enabled skip behavior --- .../templates/deployment.yaml | 4 +- .../templates/{clusterrole.yaml => role.yaml} | 6 --- ...usterrolebinding.yaml => rolebinding.yaml} | 10 ----- charts/openfga-operator/values.yaml | 5 +-- charts/openfga/templates/deployment.yaml | 1 + docs/adr/004-operator-deployment-model.md | 19 ++++---- operator/README.md | 8 ++-- operator/cmd/main.go | 20 ++++----- operator/internal/controller/helpers.go | 6 ++- .../controller/migration_controller.go | 8 +++- .../controller/migration_controller_test.go | 44 ++++++++++++++++++- 11 files changed, 79 insertions(+), 52 deletions(-) rename charts/openfga-operator/templates/{clusterrole.yaml => role.yaml} (86%) rename charts/openfga-operator/templates/{clusterrolebinding.yaml => rolebinding.yaml} (69%) diff --git a/charts/openfga-operator/templates/deployment.yaml b/charts/openfga-operator/templates/deployment.yaml index 45706601..ecad0916 100644 --- a/charts/openfga-operator/templates/deployment.yaml +++ b/charts/openfga-operator/templates/deployment.yaml @@ -40,9 +40,7 @@ spec: {{- if .Values.leaderElection.enabled }} - --leader-elect {{- end }} - {{- if .Values.watchAllNamespaces }} - - --watch-all-namespaces - {{- else if .Values.watchNamespace }} + {{- if .Values.watchNamespace }} - --watch-namespace={{ .Values.watchNamespace }} {{- end }} env: diff --git a/charts/openfga-operator/templates/clusterrole.yaml b/charts/openfga-operator/templates/role.yaml similarity index 86% rename from charts/openfga-operator/templates/clusterrole.yaml rename to charts/openfga-operator/templates/role.yaml index 652b48e3..dd17870b 100644 --- a/charts/openfga-operator/templates/clusterrole.yaml +++ b/charts/openfga-operator/templates/role.yaml @@ -1,14 +1,8 @@ apiVersion: rbac.authorization.k8s.io/v1 -{{- if .Values.watchAllNamespaces }} -kind: ClusterRole -{{- else }} kind: Role -{{- end }} metadata: name: {{ include "openfga-operator.fullname" . }} - {{- if not .Values.watchAllNamespaces }} namespace: {{ include "openfga-operator.namespace" . }} - {{- end }} labels: {{- include "openfga-operator.labels" . | nindent 4 }} rules: diff --git a/charts/openfga-operator/templates/clusterrolebinding.yaml b/charts/openfga-operator/templates/rolebinding.yaml similarity index 69% rename from charts/openfga-operator/templates/clusterrolebinding.yaml rename to charts/openfga-operator/templates/rolebinding.yaml index cfca8d1a..afacb98a 100644 --- a/charts/openfga-operator/templates/clusterrolebinding.yaml +++ b/charts/openfga-operator/templates/rolebinding.yaml @@ -1,23 +1,13 @@ apiVersion: rbac.authorization.k8s.io/v1 -{{- if .Values.watchAllNamespaces }} -kind: ClusterRoleBinding -{{- else }} kind: RoleBinding -{{- end }} metadata: name: {{ include "openfga-operator.fullname" . }} - {{- if not .Values.watchAllNamespaces }} namespace: {{ include "openfga-operator.namespace" . }} - {{- end }} labels: {{- include "openfga-operator.labels" . | nindent 4 }} roleRef: apiGroup: rbac.authorization.k8s.io - {{- if .Values.watchAllNamespaces }} - kind: ClusterRole - {{- else }} kind: Role - {{- end }} name: {{ include "openfga-operator.fullname" . }} subjects: - kind: ServiceAccount diff --git a/charts/openfga-operator/values.yaml b/charts/openfga-operator/values.yaml index 59c9dfba..38a4bff6 100644 --- a/charts/openfga-operator/values.yaml +++ b/charts/openfga-operator/values.yaml @@ -34,13 +34,10 @@ securityContext: runAsNonRoot: true runAsUser: 65532 -# -- Constrain the operator to watch a single namespace. +# -- Namespace to watch for OpenFGA Deployments. # Leave empty to default to the release namespace. watchNamespace: "" -# -- Watch all namespaces. Overrides watchNamespace. -watchAllNamespaces: false - leaderElection: # -- Enable leader election for controller manager. enabled: true diff --git a/charts/openfga/templates/deployment.yaml b/charts/openfga/templates/deployment.yaml index 2de84cdb..5fc980f5 100644 --- a/charts/openfga/templates/deployment.yaml +++ b/charts/openfga/templates/deployment.yaml @@ -6,6 +6,7 @@ metadata: {{- include "openfga.labels" . | nindent 4 }} annotations: {{- if and .Values.operator.enabled .Values.migration.enabled }} + openfga.dev/migration-enabled: "true" openfga.dev/desired-replicas: "{{ ternary 1 .Values.replicaCount (eq .Values.datastore.engine "memory") }}" {{- if or .Values.migration.serviceAccount.create .Values.migration.serviceAccount.name }} openfga.dev/migration-service-account: "{{ include "openfga.migrationServiceAccountName" . }}" diff --git a/docs/adr/004-operator-deployment-model.md b/docs/adr/004-operator-deployment-model.md index 745e7775..603bf3a2 100644 --- a/docs/adr/004-operator-deployment-model.md +++ b/docs/adr/004-operator-deployment-model.md @@ -96,10 +96,14 @@ helm-charts/ dependencies: - name: openfga-operator version: "0.1.x" - repository: "oci://ghcr.io/openfga/helm-charts" + repository: "file://../openfga-operator" condition: operator.enabled ``` +> **Note:** The `file://` reference is used because the operator subchart lives in the same +> monorepo. When the charts are published, consumers pulling from a registry will resolve the +> dependency automatically via the chart's packaging. + ### CRD Handling Helm has specific behavior around CRDs: @@ -130,19 +134,12 @@ kubectl apply -f https://github.com/openfga/helm-charts/releases/download/v0.2.0 ### Multi-Instance Considerations -When multiple OpenFGA installations exist in the same cluster: - -- **All-in-one mode:** Each installation gets its own operator instance. The operator only watches resources in its own namespace. This is simple but wasteful. -- **Standalone mode:** One operator installation watches all namespaces (or a configured set). Individual OpenFGA installations set `operator.enabled=false`. This is more efficient for large clusters. - -The operator will support both modes via a `watchNamespace` configuration: +When multiple OpenFGA installations exist in the same cluster, each installation gets its own operator instance. The operator is **namespace-scoped** — it only watches resources in its own namespace (or the namespace specified via `--watch-namespace`). This ensures independent OpenFGA installations never interfere with each other. ```yaml # Operator values operator: - watchNamespace: "" # empty = watch own namespace only (all-in-one mode) - # watchNamespace: "" # or set to a specific namespace - # watchAllNamespaces: true # watch all namespaces (standalone mode) + watchNamespace: "" # empty = watch own namespace only (default) ``` ## Consequences @@ -153,7 +150,7 @@ operator: - **Opt-out available** — `operator.enabled: false` for users who manage it separately or don't need it - **Independent versioning** — operator chart has its own version; can be released on a different cadence than the main chart - **Clean code separation** — operator code and templates are in their own chart directory -- **Standalone installation supported** — cluster admins can install one operator for multiple OpenFGA instances +- **Namespace isolation** — each operator instance is scoped to its own namespace, so multiple OpenFGA installations coexist safely - **Consistent with ecosystem** — this is the same pattern used by charts that depend on Bitnami PostgreSQL, Redis, etc. ### Negative diff --git a/operator/README.md b/operator/README.md index a2e927b1..0e43a3d1 100644 --- a/operator/README.md +++ b/operator/README.md @@ -6,7 +6,7 @@ This is **Stage 1** of the operator — focused solely on migration orchestratio ## How It Works -1. The operator watches Deployments labeled `app.kubernetes.io/part-of: openfga` and `app.kubernetes.io/component: authorization-controller` +1. The operator watches Deployments **in its own namespace** labeled `app.kubernetes.io/part-of: openfga` and `app.kubernetes.io/component: authorization-controller` 2. When a version change is detected (comparing the container image tag to the `{name}-migration-status` ConfigMap), the operator: - Keeps the Deployment at 0 replicas - Creates a migration Job running `openfga migrate` @@ -17,7 +17,7 @@ This is **Stage 1** of the operator — focused solely on migration orchestratio ## Prerequisites -- Go 1.25+ +- Go 1.26.2+ - Docker - Helm 3.6+ - A Kubernetes cluster (Rancher Desktop, kind, etc.) @@ -109,8 +109,7 @@ The operator accepts the following flags: | Flag | Default | Description | |------|---------|-------------| | `--leader-elect` | `false` | Enable leader election so only one replica actively reconciles at a time. Required when running multiple operator replicas for high availability; standby pods wait for the leader's Lease to expire before taking over. Not needed for single-replica deployments. | -| `--watch-namespace` | `""` | Namespace to watch for OpenFGA Deployments. Defaults to the operator pod's own namespace (via `POD_NAMESPACE` env var). Set explicitly for multi-namespace setups. | -| `--watch-all-namespaces` | `false` | Watch all namespaces for OpenFGA Deployments, making the operator cluster-wide. Overrides `--watch-namespace`. | +| `--watch-namespace` | `""` | Namespace to watch for OpenFGA Deployments. Defaults to the operator pod's own namespace (via `POD_NAMESPACE` env var). Each operator instance manages only its own namespace, so multiple independent OpenFGA installations can coexist safely. | | `--metrics-bind-address` | `:8080` | Address the Prometheus metrics endpoint binds to. Change only if the default port conflicts with other containers in the pod. | | `--health-probe-bind-address` | `:8081` | Address the Kubernetes liveness and readiness probe endpoints bind to. Change only if the default port conflicts. | | `--backoff-limit` | `3` | Number of times a migration Job's pod can fail before the Job is considered failed. After hitting this limit the operator deletes the Job, sets a `MigrationFailed` condition on the Deployment, and retries after a 60-second cooldown. | @@ -125,6 +124,7 @@ The operator reads these annotations from the OpenFGA Deployment: | Annotation | Description | |------------|-------------| +| `openfga.dev/migration-enabled` | Must be `"true"` for the operator to manage migrations. Deployments without this annotation are ignored. Set by the Helm chart when `operator.enabled` and `migration.enabled` are both true. | | `openfga.dev/desired-replicas` | The replica count to restore after migration succeeds. Set by the Helm chart. | | `openfga.dev/migration-service-account` | The ServiceAccount to use for migration Jobs. Defaults to the Deployment's SA. | diff --git a/operator/cmd/main.go b/operator/cmd/main.go index 9fb91070..79172649 100644 --- a/operator/cmd/main.go +++ b/operator/cmd/main.go @@ -23,19 +23,17 @@ func init() { func main() { var ( - leaderElect bool - watchNamespace string - watchAllNamespaces bool - metricsAddr string - healthProbeAddr string - backoffLimit int - activeDeadline int - ttlAfterFinished int + leaderElect bool + watchNamespace string + metricsAddr string + healthProbeAddr string + backoffLimit int + activeDeadline int + ttlAfterFinished int ) flag.BoolVar(&leaderElect, "leader-elect", false, "Enable leader election for the controller manager.") flag.StringVar(&watchNamespace, "watch-namespace", "", "Namespace to watch. Defaults to the operator pod namespace.") - flag.BoolVar(&watchAllNamespaces, "watch-all-namespaces", false, "Watch all namespaces.") flag.StringVar(&metricsAddr, "metrics-bind-address", ":8080", "The address the metric endpoint binds to.") flag.StringVar(&healthProbeAddr, "health-probe-bind-address", ":8081", "The address the health probe endpoint binds to.") flag.IntVar(&backoffLimit, "backoff-limit", int(controller.DefaultBackoffLimit), "BackoffLimit for migration Jobs.") @@ -50,7 +48,7 @@ func main() { logger := ctrl.Log.WithName("setup") // Fall back to the pod's namespace when no explicit scope is set. - if !watchAllNamespaces && watchNamespace == "" { + if watchNamespace == "" { if podNS, ok := os.LookupEnv("POD_NAMESPACE"); ok && podNS != "" { watchNamespace = podNS logger.Info("defaulting watch scope to pod namespace", "namespace", podNS) @@ -59,7 +57,7 @@ func main() { // Configure cache namespace restrictions. var cacheOpts cache.Options - if watchNamespace != "" && !watchAllNamespaces { + if watchNamespace != "" { cacheOpts.DefaultNamespaces = map[string]cache.Config{ watchNamespace: {}, } diff --git a/operator/internal/controller/helpers.go b/operator/internal/controller/helpers.go index a9c37d51..efc7a4ba 100644 --- a/operator/internal/controller/helpers.go +++ b/operator/internal/controller/helpers.go @@ -24,6 +24,7 @@ const ( LabelComponentValue = "authorization-controller" // Annotations set on the Deployment by the Helm chart / operator. + AnnotationMigrationEnabled = "openfga.dev/migration-enabled" AnnotationDesiredReplicas = "openfga.dev/desired-replicas" AnnotationMigrationServiceAccount = "openfga.dev/migration-service-account" AnnotationRetryAfter = "openfga.dev/migration-retry-after" @@ -150,10 +151,13 @@ func buildMigrationJob( Image: mainContainer.Image, Args: []string{"migrate"}, Env: datastoreEnvVars, + EnvFrom: mainContainer.EnvFrom, + VolumeMounts: mainContainer.VolumeMounts, SecurityContext: mainContainer.SecurityContext, }, }, - // Inherit scheduling constraints from the parent Deployment. + // Inherit volumes and scheduling constraints from the parent Deployment. + Volumes: deployment.Spec.Template.Spec.Volumes, NodeSelector: deployment.Spec.Template.Spec.NodeSelector, Tolerations: deployment.Spec.Template.Spec.Tolerations, Affinity: deployment.Spec.Template.Spec.Affinity, diff --git a/operator/internal/controller/migration_controller.go b/operator/internal/controller/migration_controller.go index 77dd22eb..e28aed86 100644 --- a/operator/internal/controller/migration_controller.go +++ b/operator/internal/controller/migration_controller.go @@ -47,7 +47,13 @@ func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) ( return ctrl.Result{}, err } - // 2. Find the OpenFGA container and extract the desired version. + // 2. Skip if migration is not opted-in via annotation. + if deployment.Annotations[AnnotationMigrationEnabled] != "true" { + logger.V(1).Info("migration not enabled for this deployment, skipping") + return ctrl.Result{}, nil + } + + // 3. Find the OpenFGA container and extract the desired version. mainContainer := findOpenFGAContainer(deployment) if mainContainer == nil { logger.Info("deployment has no containers, skipping") diff --git a/operator/internal/controller/migration_controller_test.go b/operator/internal/controller/migration_controller_test.go index f463c7d9..811e89c8 100644 --- a/operator/internal/controller/migration_controller_test.go +++ b/operator/internal/controller/migration_controller_test.go @@ -33,7 +33,9 @@ func newTestDeployment(name, namespace, image string, replicas int32) *appsv1.De LabelPartOf: LabelPartOfValue, LabelComponent: LabelComponentValue, }, - Annotations: map[string]string{}, + Annotations: map[string]string{ + AnnotationMigrationEnabled: "true", + }, }, Spec: appsv1.DeploymentSpec{ Replicas: ptr.To(replicas), @@ -438,6 +440,46 @@ func TestReconcile_FindContainerByName(t *testing.T) { } } +func TestReconcile_MigrationNotEnabled_Skips(t *testing.T) { + // Given: a Deployment without the migration-enabled annotation. + dep := newTestDeployment("openfga", "default", "openfga/openfga:v1.14.0", 3) + delete(dep.Annotations, AnnotationMigrationEnabled) + + r := newReconciler(dep) + + // When: reconciling. + result, err := r.Reconcile(context.Background(), ctrl.Request{ + NamespacedName: types.NamespacedName{Name: "openfga", Namespace: "default"}, + }) + + // Then: no error, no requeue, no Job created, replicas unchanged. + if err != nil { + t.Fatalf("unexpected error: %v", err) + } + if result.RequeueAfter != 0 { + t.Error("expected no requeue when migration is not enabled") + } + + // Verify no Job was created. + job := &batchv1.Job{} + if getErr := r.Get(context.Background(), types.NamespacedName{ + Name: "openfga-migrate", Namespace: "default", + }, job); getErr == nil { + t.Error("expected no migration job when migration is not enabled") + } + + // Verify replicas unchanged. + updated := &appsv1.Deployment{} + if getErr := r.Get(context.Background(), types.NamespacedName{ + Name: "openfga", Namespace: "default", + }, updated); getErr != nil { + t.Fatalf("getting deployment: %v", getErr) + } + if *updated.Spec.Replicas != 3 { + t.Errorf("expected 3 replicas unchanged, got %d", *updated.Spec.Replicas) + } +} + func TestExtractImageTag(t *testing.T) { tests := []struct { image string From 8c725030b3a4793301e41f140408b868e3e3bd5a Mon Sep 17 00:00:00 2001 From: Ed Milic <edmilic@gmail.com> Date: Sat, 11 Apr 2026 10:12:47 -0400 Subject: [PATCH 06/70] fix: remove EnvFrom from migration Job to preserve least-privilege The migration Job should only receive explicitly filtered OPENFGA_DATASTORE_* env vars, not the full EnvFrom from the source Deployment which could leak non-datastore secrets. --- operator/internal/controller/helpers.go | 1 - 1 file changed, 1 deletion(-) diff --git a/operator/internal/controller/helpers.go b/operator/internal/controller/helpers.go index efc7a4ba..3727e2f0 100644 --- a/operator/internal/controller/helpers.go +++ b/operator/internal/controller/helpers.go @@ -151,7 +151,6 @@ func buildMigrationJob( Image: mainContainer.Image, Args: []string{"migrate"}, Env: datastoreEnvVars, - EnvFrom: mainContainer.EnvFrom, VolumeMounts: mainContainer.VolumeMounts, SecurityContext: mainContainer.SecurityContext, }, From 8c857ac52a0e3751b4892848257a372c23d12dda Mon Sep 17 00:00:00 2001 From: Ed Milic <edmilic@gmail.com> Date: Sat, 11 Apr 2026 10:24:29 -0400 Subject: [PATCH 07/70] fix: address Copilot review round 3 on PR #309 - Wrap deployment annotations in conditional to avoid emitting empty annotations: field which produces an invalid manifest - Store full version in annotation (openfga.dev/desired-version) and truncate label to 63 chars to support digest-pinned images - Align operator image default to ghcr.io/openfga/openfga-operator to match CI publishing target --- .github/workflows/test.yml | 4 ++-- charts/openfga-operator/values.yaml | 2 +- charts/openfga/templates/deployment.yaml | 5 ++++- operator/internal/controller/helpers.go | 11 ++++++++++- operator/internal/controller/migration_controller.go | 7 ++++++- 5 files changed, 23 insertions(+), 6 deletions(-) diff --git a/.github/workflows/test.yml b/.github/workflows/test.yml index f2369c54..bbec31d6 100644 --- a/.github/workflows/test.yml +++ b/.github/workflows/test.yml @@ -63,8 +63,8 @@ jobs: if: steps.list-changed.outputs.changed == 'true' run: | version=$(grep '^appVersion:' charts/openfga-operator/Chart.yaml | awk '{print $2}' | tr -d '"') - docker build -t "openfga/openfga-operator:${version}" operator/ - kind load docker-image "openfga/openfga-operator:${version}" --name chart-testing + docker build -t "ghcr.io/openfga/openfga-operator:${version}" operator/ + kind load docker-image "ghcr.io/openfga/openfga-operator:${version}" --name chart-testing - name: Run chart-testing (install) if: steps.list-changed.outputs.changed == 'true' diff --git a/charts/openfga-operator/values.yaml b/charts/openfga-operator/values.yaml index 38a4bff6..1b525258 100644 --- a/charts/openfga-operator/values.yaml +++ b/charts/openfga-operator/values.yaml @@ -1,7 +1,7 @@ replicaCount: 1 image: - repository: openfga/openfga-operator + repository: ghcr.io/openfga/openfga-operator pullPolicy: IfNotPresent # -- Overrides the image tag whose default is the chart appVersion. tag: "" diff --git a/charts/openfga/templates/deployment.yaml b/charts/openfga/templates/deployment.yaml index 5fc980f5..ded90abb 100644 --- a/charts/openfga/templates/deployment.yaml +++ b/charts/openfga/templates/deployment.yaml @@ -4,8 +4,10 @@ metadata: name: {{ include "openfga.fullname" . }} labels: {{- include "openfga.labels" . | nindent 4 }} + {{- $hasOperatorAnnotations := and .Values.operator.enabled .Values.migration.enabled }} + {{- if or $hasOperatorAnnotations .Values.annotations }} annotations: - {{- if and .Values.operator.enabled .Values.migration.enabled }} + {{- if $hasOperatorAnnotations }} openfga.dev/migration-enabled: "true" openfga.dev/desired-replicas: "{{ ternary 1 .Values.replicaCount (eq .Values.datastore.engine "memory") }}" {{- if or .Values.migration.serviceAccount.create .Values.migration.serviceAccount.name }} @@ -15,6 +17,7 @@ metadata: {{- with .Values.annotations }} {{- toYaml . | nindent 4 }} {{- end }} + {{- end }} spec: {{- if and .Values.operator.enabled .Values.migration.enabled }} {{- if .Values.autoscaling.enabled }} diff --git a/operator/internal/controller/helpers.go b/operator/internal/controller/helpers.go index 3727e2f0..799c6099 100644 --- a/operator/internal/controller/helpers.go +++ b/operator/internal/controller/helpers.go @@ -108,6 +108,12 @@ func buildMigrationJob( } } + // Truncate version for label (max 63 chars); store full version in annotation. + labelVersion := desiredVersion + if len(labelVersion) > 63 { + labelVersion = labelVersion[:63] + } + return &batchv1.Job{ ObjectMeta: metav1.ObjectMeta{ Name: migrationJobName(deployment.Name), @@ -116,7 +122,10 @@ func buildMigrationJob( LabelPartOf: LabelPartOfValue, LabelComponent: "migration", "app.kubernetes.io/managed-by": "openfga-operator", - "app.kubernetes.io/version": desiredVersion, + "app.kubernetes.io/version": labelVersion, + }, + Annotations: map[string]string{ + "openfga.dev/desired-version": desiredVersion, }, OwnerReferences: []metav1.OwnerReference{ { diff --git a/operator/internal/controller/migration_controller.go b/operator/internal/controller/migration_controller.go index e28aed86..0e873250 100644 --- a/operator/internal/controller/migration_controller.go +++ b/operator/internal/controller/migration_controller.go @@ -146,7 +146,12 @@ func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) ( } // 8b. If the existing Job is for a different version, delete it and recreate. - if jobVersion := job.Labels["app.kubernetes.io/version"]; jobVersion != "" && jobVersion != desiredVersion { + // Check annotation first (supports digests > 63 chars), fall back to label. + jobVersion := job.Annotations["openfga.dev/desired-version"] + if jobVersion == "" { + jobVersion = job.Labels["app.kubernetes.io/version"] + } + if jobVersion != "" && jobVersion != desiredVersion { logger.Info("existing migration job is for a different version, deleting", "jobVersion", jobVersion, "desiredVersion", desiredVersion) propagation := metav1.DeletePropagationBackground if delErr := r.Delete(ctx, job, &client.DeleteOptions{ From 0cdaac8723256ee67d9e9a8cd9d410f50ec47200 Mon Sep 17 00:00:00 2001 From: Ed Milic <edmilic@gmail.com> Date: Sat, 11 Apr 2026 11:48:59 -0400 Subject: [PATCH 08/70] fix: desired version to replace problematic ":" with "_" for label value" --- operator/internal/controller/helpers.go | 5 +++-- 1 file changed, 3 insertions(+), 2 deletions(-) diff --git a/operator/internal/controller/helpers.go b/operator/internal/controller/helpers.go index 799c6099..871d1a55 100644 --- a/operator/internal/controller/helpers.go +++ b/operator/internal/controller/helpers.go @@ -108,8 +108,9 @@ func buildMigrationJob( } } - // Truncate version for label (max 63 chars); store full version in annotation. - labelVersion := desiredVersion + // Sanitize version for use as a label value (must match [a-zA-Z0-9._-], max 63 chars). + // The full version is stored in an annotation for accurate comparison. + labelVersion := strings.ReplaceAll(desiredVersion, ":", "_") if len(labelVersion) > 63 { labelVersion = labelVersion[:63] } From a7c84460325c32ee8f08714b08fa5a0913f6caae Mon Sep 17 00:00:00 2001 From: Ed Milic <edmilic@gmail.com> Date: Sat, 11 Apr 2026 12:04:42 -0400 Subject: [PATCH 09/70] fix: address Copilot review comments - Return error on retry-after annotation patch failure to prevent Job churn that bypasses the 60s cooldown - Add test for stale-Job version mismatch deletion path --- .../controller/migration_controller.go | 2 +- .../controller/migration_controller_test.go | 72 +++++++++++++++++++ 2 files changed, 73 insertions(+), 1 deletion(-) diff --git a/operator/internal/controller/migration_controller.go b/operator/internal/controller/migration_controller.go index 0e873250..1cf0f06b 100644 --- a/operator/internal/controller/migration_controller.go +++ b/operator/internal/controller/migration_controller.go @@ -210,7 +210,7 @@ func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) ( } deployment.Annotations[AnnotationRetryAfter] = retryAfter if patchErr := r.Patch(ctx, deployment, patch); patchErr != nil { - logger.Error(patchErr, "failed to set retry-after annotation") + return ctrl.Result{}, fmt.Errorf("persisting retry-after annotation: %w", patchErr) } // Delete the failed Job so a fresh one is created on the next reconcile. diff --git a/operator/internal/controller/migration_controller_test.go b/operator/internal/controller/migration_controller_test.go index 811e89c8..636182c7 100644 --- a/operator/internal/controller/migration_controller_test.go +++ b/operator/internal/controller/migration_controller_test.go @@ -440,6 +440,78 @@ func TestReconcile_FindContainerByName(t *testing.T) { } } +func TestReconcile_StaleJob_DeletedAndRequeued(t *testing.T) { + // Given: a Deployment at v1.15.0 with an existing migration Job for v1.14.0. + dep := newTestDeployment("openfga", "default", "openfga/openfga:v1.15.0", 0) + dep.Annotations[AnnotationDesiredReplicas] = "3" + + staleJob := &batchv1.Job{ + ObjectMeta: metav1.ObjectMeta{ + Name: "openfga-migrate", + Namespace: "default", + Labels: map[string]string{ + "app.kubernetes.io/version": "v1.14.0", + }, + Annotations: map[string]string{ + "openfga.dev/desired-version": "v1.14.0", + }, + OwnerReferences: []metav1.OwnerReference{ + { + APIVersion: "apps/v1", + Kind: "Deployment", + Name: "openfga", + UID: "test-uid-123", + }, + }, + }, + Spec: batchv1.JobSpec{ + BackoffLimit: ptr.To(int32(3)), + Template: corev1.PodTemplateSpec{ + Spec: corev1.PodSpec{ + Containers: []corev1.Container{{Name: "migrate", Image: "openfga/openfga:v1.14.0"}}, + RestartPolicy: corev1.RestartPolicyNever, + }, + }, + }, + Status: batchv1.JobStatus{ + Succeeded: 1, + }, + } + + r := newReconciler(dep, staleJob) + + // When: reconciling. + result, err := r.Reconcile(context.Background(), ctrl.Request{ + NamespacedName: types.NamespacedName{Name: "openfga", Namespace: "default"}, + }) + + // Then: no error, requeue to recreate with correct version. + if err != nil { + t.Fatalf("unexpected error: %v", err) + } + if result.RequeueAfter == 0 { + t.Error("expected requeue after deleting stale job") + } + + // Verify the stale Job was deleted. + deletedJob := &batchv1.Job{} + if getErr := r.Get(context.Background(), types.NamespacedName{ + Name: "openfga-migrate", Namespace: "default", + }, deletedJob); getErr == nil { + t.Error("expected stale migration job to be deleted") + } + + // Verify ConfigMap was NOT updated (migration didn't actually run for v1.15.0). + cm := &corev1.ConfigMap{} + if getErr := r.Get(context.Background(), types.NamespacedName{ + Name: "openfga-migration-status", Namespace: "default", + }, cm); getErr == nil { + if cm.Data["version"] == "v1.15.0" { + t.Error("ConfigMap should not be updated to v1.15.0 from a stale v1.14.0 job") + } + } +} + func TestReconcile_MigrationNotEnabled_Skips(t *testing.T) { // Given: a Deployment without the migration-enabled annotation. dep := newTestDeployment("openfga", "default", "openfga/openfga:v1.14.0", 3) From 95f30e76339fdd5f00485989f239e5a2e3e5fd45 Mon Sep 17 00:00:00 2001 From: Ed Milic <edmilic@gmail.com> Date: Sat, 11 Apr 2026 12:30:15 -0400 Subject: [PATCH 10/70] fix: validate flags to prevent negative or out-of-range values (i.e. values overflowing) --- operator/cmd/main.go | 18 ++++++++++++++++++ 1 file changed, 18 insertions(+) diff --git a/operator/cmd/main.go b/operator/cmd/main.go index 79172649..ac9bac64 100644 --- a/operator/cmd/main.go +++ b/operator/cmd/main.go @@ -2,6 +2,8 @@ package main import ( "flag" + "fmt" + "math" "os" "k8s.io/apimachinery/pkg/runtime" @@ -44,6 +46,22 @@ func main() { opts.BindFlags(flag.CommandLine) flag.Parse() + // Validate flag values. + for _, v := range []struct { + name string + value int + max int + }{ + {"backoff-limit", backoffLimit, math.MaxInt32}, + {"active-deadline-seconds", activeDeadline, math.MaxInt32}, + {"ttl-seconds-after-finished", ttlAfterFinished, math.MaxInt32}, + } { + if v.value < 0 || v.value > v.max { + fmt.Fprintf(os.Stderr, "invalid value for --%s: must be between 0 and %d\n", v.name, v.max) + os.Exit(1) + } + } + ctrl.SetLogger(zap.New(zap.UseFlagOptions(&opts))) logger := ctrl.Log.WithName("setup") From 5545a95b5410ce8d283fca4c510a09f6921855ce Mon Sep 17 00:00:00 2001 From: Ed Milic <edmilic@gmail.com> Date: Sat, 11 Apr 2026 12:37:17 -0400 Subject: [PATCH 11/70] fix: gate legacy migration initContainers on operator.enabled When operator.enabled=true but migration.enabled=false, the legacy wait-for-migration initContainer could render and hang waiting for a Job that will never be created. Gate on operator.enabled instead so the legacy path is fully disabled when the operator is installed. --- charts/openfga/templates/deployment.yaml | 3 +-- 1 file changed, 1 insertion(+), 2 deletions(-) diff --git a/charts/openfga/templates/deployment.yaml b/charts/openfga/templates/deployment.yaml index ded90abb..275bd869 100644 --- a/charts/openfga/templates/deployment.yaml +++ b/charts/openfga/templates/deployment.yaml @@ -52,8 +52,7 @@ spec: serviceAccountName: {{ include "openfga.serviceAccountName" . }} securityContext: {{- toYaml .Values.podSecurityContext | nindent 8 }} - {{- $operatorMigration := and .Values.operator.enabled .Values.migration.enabled }} - {{- $needsMigrationInit := and (not $operatorMigration) (has .Values.datastore.engine (list "postgres" "mysql")) .Values.datastore.applyMigrations .Values.datastore.waitForMigrations }} + {{- $needsMigrationInit := and (not .Values.operator.enabled) (has .Values.datastore.engine (list "postgres" "mysql")) .Values.datastore.applyMigrations .Values.datastore.waitForMigrations }} {{- if or $needsMigrationInit .Values.extraInitContainers }} initContainers: {{- if $needsMigrationInit }} From 80a55fa888aed6bc4cfadbeb87b9bfd4aeea8897 Mon Sep 17 00:00:00 2001 From: Ed Milic <edmilic@gmail.com> Date: Sat, 11 Apr 2026 12:53:39 -0400 Subject: [PATCH 12/70] fix: use single quotes for Helm annotations with nested template expressions The double-quoted annotation values contained inner double quotes (e.g. "memory") which produced invalid YAML that IDEs flagged as errors. Switch to single-quote wrappers so the inner Go template strings don't conflict. --- charts/openfga/templates/deployment.yaml | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/charts/openfga/templates/deployment.yaml b/charts/openfga/templates/deployment.yaml index 275bd869..a70c93dc 100644 --- a/charts/openfga/templates/deployment.yaml +++ b/charts/openfga/templates/deployment.yaml @@ -9,9 +9,9 @@ metadata: annotations: {{- if $hasOperatorAnnotations }} openfga.dev/migration-enabled: "true" - openfga.dev/desired-replicas: "{{ ternary 1 .Values.replicaCount (eq .Values.datastore.engine "memory") }}" + openfga.dev/desired-replicas: '{{ ternary 1 .Values.replicaCount (eq .Values.datastore.engine "memory") }}' {{- if or .Values.migration.serviceAccount.create .Values.migration.serviceAccount.name }} - openfga.dev/migration-service-account: "{{ include "openfga.migrationServiceAccountName" . }}" + openfga.dev/migration-service-account: '{{ include "openfga.migrationServiceAccountName" . }}' {{- end }} {{- end }} {{- with .Values.annotations }} From af2ff38a42d4fee6564fdacf2d2d35cce4661d17 Mon Sep 17 00:00:00 2001 From: Ed Milic <edmilic@gmail.com> Date: Sat, 11 Apr 2026 12:54:47 -0400 Subject: [PATCH 13/70] fix: address Copilot review round 6 on PR #309 - Pin Dockerfile base images by digest for reproducible builds - Fix version label fallback comparison for digest-pinned images by sanitizing desiredVersion before comparing to the label value --- operator/Dockerfile | 6 ++++-- operator/internal/controller/migration_controller.go | 9 ++++++++- 2 files changed, 12 insertions(+), 3 deletions(-) diff --git a/operator/Dockerfile b/operator/Dockerfile index 7d836a3a..846414a1 100644 --- a/operator/Dockerfile +++ b/operator/Dockerfile @@ -1,4 +1,5 @@ -FROM golang:1.26.2 AS builder +# pinned golang:1.26.2 linux/amd64 +FROM golang:1.26.2@sha256:b53c282df83967299380adbd6a2dc67e750a58217f39285d6240f6f80b19eaad AS builder WORKDIR /workspace COPY go.mod go.sum ./ @@ -9,7 +10,8 @@ COPY internal/ internal/ RUN CGO_ENABLED=0 GOOS=linux go build -ldflags="-s -w" -o /operator ./cmd/ -FROM gcr.io/distroless/static:nonroot +# pinned gcr.io/distroless/static:nonroot linux/amd64 +FROM gcr.io/distroless/static:nonroot@sha256:64c43684e6d2b581d1eb362ea47b6a4defee6a9cac5f7ebbda3daa67e8c9b8e6 WORKDIR / COPY --from=builder /operator . USER 65532:65532 diff --git a/operator/internal/controller/migration_controller.go b/operator/internal/controller/migration_controller.go index 1cf0f06b..be00b7ce 100644 --- a/operator/internal/controller/migration_controller.go +++ b/operator/internal/controller/migration_controller.go @@ -148,10 +148,17 @@ func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) ( // 8b. If the existing Job is for a different version, delete it and recreate. // Check annotation first (supports digests > 63 chars), fall back to label. jobVersion := job.Annotations["openfga.dev/desired-version"] + versionMatch := jobVersion == desiredVersion if jobVersion == "" { + // Label values have ":" replaced with "_", so sanitize desiredVersion for comparison. + sanitized := strings.ReplaceAll(desiredVersion, ":", "_") + if len(sanitized) > 63 { + sanitized = sanitized[:63] + } jobVersion = job.Labels["app.kubernetes.io/version"] + versionMatch = jobVersion == sanitized } - if jobVersion != "" && jobVersion != desiredVersion { + if jobVersion != "" && !versionMatch { logger.Info("existing migration job is for a different version, deleting", "jobVersion", jobVersion, "desiredVersion", desiredVersion) propagation := metav1.DeletePropagationBackground if delErr := r.Delete(ctx, job, &client.DeleteOptions{ From b29322d4751006c722d746dec5d51fc595896757 Mon Sep 17 00:00:00 2001 From: Ed Milic <edmilic@gmail.com> Date: Sat, 11 Apr 2026 14:13:27 -0400 Subject: [PATCH 14/70] fix: remove env var filtering, fix tests, updated README.md to clarify migration-specific volumes --- operator/README.md | 1 + operator/internal/controller/helpers.go | 11 +--- .../controller/migration_controller_test.go | 57 +++++++++++++++++-- 3 files changed, 55 insertions(+), 14 deletions(-) diff --git a/operator/README.md b/operator/README.md index 0e43a3d1..c9efbe8e 100644 --- a/operator/README.md +++ b/operator/README.md @@ -131,3 +131,4 @@ The operator reads these annotations from the OpenFGA Deployment: ## Limitations - **Mutable image tags:** The operator detects version changes by comparing the container image tag (or digest). If you deploy with a mutable tag like `latest` or reuse the same tag for different builds, the operator will not detect changes and will skip the migration. Use immutable tags (e.g., `v1.14.0`) or pin images by digest for reliable migration triggering. +- **Migration-specific volumes:** The legacy Helm chart values `migrate.extraVolumes` and `migrate.extraVolumeMounts` have no effect in operator mode. The operator inherits volumes and mounts from the main Deployment pod spec. If you need additional volumes for migrations (e.g., CA bundles or TLS certs), add them to the top-level `extraVolumes` and `extraVolumeMounts` values instead. diff --git a/operator/internal/controller/helpers.go b/operator/internal/controller/helpers.go index 871d1a55..deaa9e1e 100644 --- a/operator/internal/controller/helpers.go +++ b/operator/internal/controller/helpers.go @@ -100,14 +100,6 @@ func buildMigrationJob( migrationSA = deployment.Spec.Template.Spec.ServiceAccountName } - // Filter env vars — only pass datastore-related vars to the migration Job. - var datastoreEnvVars []corev1.EnvVar - for _, env := range mainContainer.Env { - if strings.HasPrefix(env.Name, "OPENFGA_DATASTORE_") { - datastoreEnvVars = append(datastoreEnvVars, env) - } - } - // Sanitize version for use as a label value (must match [a-zA-Z0-9._-], max 63 chars). // The full version is stored in an annotation for accurate comparison. labelVersion := strings.ReplaceAll(desiredVersion, ":", "_") @@ -160,7 +152,8 @@ func buildMigrationJob( Name: "migrate-database", Image: mainContainer.Image, Args: []string{"migrate"}, - Env: datastoreEnvVars, + Env: mainContainer.Env, + EnvFrom: mainContainer.EnvFrom, VolumeMounts: mainContainer.VolumeMounts, SecurityContext: mainContainer.SecurityContext, }, diff --git a/operator/internal/controller/migration_controller_test.go b/operator/internal/controller/migration_controller_test.go index 636182c7..be37fe3c 100644 --- a/operator/internal/controller/migration_controller_test.go +++ b/operator/internal/controller/migration_controller_test.go @@ -67,7 +67,8 @@ func newTestDeployment(name, namespace, image string, replicas int32) *appsv1.De func newReconciler(objects ...runtime.Object) *MigrationReconciler { scheme := newScheme() - clientBuilder := fake.NewClientBuilder().WithScheme(scheme) + clientBuilder := fake.NewClientBuilder().WithScheme(scheme). + WithStatusSubresource(&appsv1.Deployment{}) for _, obj := range objects { clientBuilder = clientBuilder.WithRuntimeObjects(obj) } @@ -79,6 +80,15 @@ func newReconciler(objects ...runtime.Object) *MigrationReconciler { } } +func findCondition(conditions []appsv1.DeploymentCondition, condType string) *appsv1.DeploymentCondition { + for i := range conditions { + if string(conditions[i].Type) == condType { + return &conditions[i] + } + } + return nil +} + func TestReconcile_FirstInstall_CreatesJob(t *testing.T) { // Given: a Deployment with no migration-status ConfigMap. dep := newTestDeployment("openfga", "default", "openfga/openfga:v1.14.0", 0) @@ -113,10 +123,14 @@ func TestReconcile_FirstInstall_CreatesJob(t *testing.T) { t.Errorf("expected job args [migrate], got %v", job.Spec.Template.Spec.Containers[0].Args) } - // Verify only datastore env vars were passed. + // Verify all env vars from the main container were passed. + jobEnvNames := make(map[string]bool) for _, env := range job.Spec.Template.Spec.Containers[0].Env { - if env.Name == "OPENFGA_LOG_LEVEL" { - t.Error("non-datastore env var OPENFGA_LOG_LEVEL should not be passed to migration job") + jobEnvNames[env.Name] = true + } + for _, expected := range []string{"OPENFGA_DATASTORE_ENGINE", "OPENFGA_DATASTORE_URI", "OPENFGA_LOG_LEVEL"} { + if !jobEnvNames[expected] { + t.Errorf("expected env var %s to be passed to migration job", expected) } } } @@ -166,9 +180,18 @@ func TestReconcile_VersionMatch_ScalesUp(t *testing.T) { } func TestReconcile_JobSucceeded_UpdatesConfigMapAndScalesUp(t *testing.T) { - // Given: a Deployment at 0 replicas, no ConfigMap, and a succeeded migration Job. + // Given: a Deployment at 0 replicas, no ConfigMap, a succeeded migration Job, + // and a pre-existing MigrationFailed condition from a prior attempt. dep := newTestDeployment("openfga", "default", "openfga/openfga:v1.14.0", 0) dep.Annotations[AnnotationDesiredReplicas] = "3" + dep.Status.Conditions = []appsv1.DeploymentCondition{ + { + Type: "MigrationFailed", + Status: corev1.ConditionTrue, + Reason: "MigrationJobFailed", + Message: "Database migration failed for version v1.13.0.", + }, + } job := &batchv1.Job{ ObjectMeta: metav1.ObjectMeta{ @@ -230,6 +253,18 @@ func TestReconcile_JobSucceeded_UpdatesConfigMapAndScalesUp(t *testing.T) { if *updated.Spec.Replicas != 3 { t.Errorf("expected 3 replicas, got %d", *updated.Spec.Replicas) } + + // Verify MigrationFailed condition was cleared. + cond := findCondition(updated.Status.Conditions, "MigrationFailed") + if cond == nil { + t.Fatal("expected MigrationFailed condition to exist") + } + if cond.Status != corev1.ConditionFalse { + t.Errorf("expected MigrationFailed status False after success, got %s", cond.Status) + } + if cond.Reason != "MigrationSucceeded" { + t.Errorf("expected reason MigrationSucceeded, got %s", cond.Reason) + } } func TestReconcile_JobFailed_SetsRetryAnnotationAndRequeues(t *testing.T) { @@ -302,6 +337,18 @@ func TestReconcile_JobFailed_SetsRetryAnnotationAndRequeues(t *testing.T) { if _, ok := updated.Annotations[AnnotationRetryAfter]; !ok { t.Error("expected retry-after annotation to be set on Deployment") } + + // Verify MigrationFailed condition was set. + cond := findCondition(updated.Status.Conditions, "MigrationFailed") + if cond == nil { + t.Fatal("expected MigrationFailed condition to be set") + } + if cond.Status != corev1.ConditionTrue { + t.Errorf("expected MigrationFailed status True, got %s", cond.Status) + } + if cond.Reason != "MigrationJobFailed" { + t.Errorf("expected reason MigrationJobFailed, got %s", cond.Reason) + } } func TestReconcile_RetryAfterCooldown_SkipsJobCreation(t *testing.T) { From d1f4ff7f962db93aa855a9ca6bb4bad9e19aa892 Mon Sep 17 00:00:00 2001 From: Ed Milic <edmilic@gmail.com> Date: Sat, 11 Apr 2026 14:31:11 -0400 Subject: [PATCH 15/70] fix: update includes OwnerReferences --- operator/internal/controller/helpers.go | 4 +++- 1 file changed, 3 insertions(+), 1 deletion(-) diff --git a/operator/internal/controller/helpers.go b/operator/internal/controller/helpers.go index deaa9e1e..7a4ac2ab 100644 --- a/operator/internal/controller/helpers.go +++ b/operator/internal/controller/helpers.go @@ -219,9 +219,11 @@ func updateMigrationStatus( return nil } - // Update existing ConfigMap. + // Update existing ConfigMap (including OwnerReferences in case the Deployment + // was deleted and recreated with a new UID). existing.Data = cm.Data existing.Labels = cm.Labels + existing.OwnerReferences = cm.OwnerReferences if updateErr := c.Update(ctx, existing); updateErr != nil { return fmt.Errorf("updating migration status ConfigMap: %w", updateErr) } From 00ba96def098278c961696e40d7854d934c9fb4a Mon Sep 17 00:00:00 2001 From: Ed Milic <edmilic@gmail.com> Date: Sat, 11 Apr 2026 14:36:25 -0400 Subject: [PATCH 16/70] feat: add PodDisruptionBudget to operator subchart Prevents the operator from being evicted during node drains, which could leave the OpenFGA Deployment stuck at 0 replicas with no controller to scale it back up. Disabled by default; supports both minAvailable and maxUnavailable modes. --- charts/openfga-operator/templates/pdb.yaml | 18 ++++++++++++++++++ charts/openfga-operator/values.yaml | 10 ++++++++++ 2 files changed, 28 insertions(+) create mode 100644 charts/openfga-operator/templates/pdb.yaml diff --git a/charts/openfga-operator/templates/pdb.yaml b/charts/openfga-operator/templates/pdb.yaml new file mode 100644 index 00000000..6c3514eb --- /dev/null +++ b/charts/openfga-operator/templates/pdb.yaml @@ -0,0 +1,18 @@ +{{- if .Values.podDisruptionBudget.enabled -}} +apiVersion: policy/v1 +kind: PodDisruptionBudget +metadata: + name: {{ include "openfga-operator.fullname" . }} + namespace: {{ include "openfga-operator.namespace" . }} + labels: + {{- include "openfga-operator.labels" . | nindent 4 }} +spec: + {{- if .Values.podDisruptionBudget.minAvailable }} + minAvailable: {{ .Values.podDisruptionBudget.minAvailable }} + {{- else }} + maxUnavailable: {{ .Values.podDisruptionBudget.maxUnavailable | default 1 }} + {{- end }} + selector: + matchLabels: + {{- include "openfga-operator.selectorLabels" . | nindent 6 }} +{{- end }} diff --git a/charts/openfga-operator/values.yaml b/charts/openfga-operator/values.yaml index 1b525258..4b308a21 100644 --- a/charts/openfga-operator/values.yaml +++ b/charts/openfga-operator/values.yaml @@ -49,6 +49,16 @@ resources: {} # limits: # memory: 128Mi +podDisruptionBudget: + # -- Enable a PodDisruptionBudget for the operator. + enabled: false + # -- Minimum number of pods that must be available during disruption. + # Cannot be set together with maxUnavailable. + minAvailable: "" + # -- Maximum number of pods that can be unavailable during disruption. + # Defaults to 1 when enabled and minAvailable is not set. + maxUnavailable: 1 + nodeSelector: {} tolerations: [] From cb2bb8a6a6750c335cb67f79d0b047ea06d04284 Mon Sep 17 00:00:00 2001 From: Ed Milic <edmilic@gmail.com> Date: Sat, 11 Apr 2026 14:38:23 -0400 Subject: [PATCH 17/70] fix: use Job conditions instead of status counters for failure detection Replace checks on job.Status.Failed >= backoffLimit with isJobConditionTrue(job, batchv1.JobFailed), and job.Status.Succeeded with batchv1.JobComplete. The Job controller sets conditions atomically when it makes its final decision, avoiding races where the operator acts on intermediate counter states before Kubernetes has finished cleaning up. --- .../controller/migration_controller.go | 23 ++++++++++++------- .../controller/migration_controller_test.go | 18 +++++++++++++++ 2 files changed, 33 insertions(+), 8 deletions(-) diff --git a/operator/internal/controller/migration_controller.go b/operator/internal/controller/migration_controller.go index be00b7ce..75cbf199 100644 --- a/operator/internal/controller/migration_controller.go +++ b/operator/internal/controller/migration_controller.go @@ -169,8 +169,8 @@ func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) ( return ctrl.Result{RequeueAfter: 5 * time.Second}, nil } - // 9. Check Job status. - if job.Status.Succeeded >= 1 { + // 9. Check Job status using conditions for authoritative completion signals. + if isJobConditionTrue(job, batchv1.JobComplete) { logger.Info("migration succeeded", "version", desiredVersion) // Clear MigrationFailed condition. @@ -193,12 +193,7 @@ func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) ( return ctrl.Result{}, nil } - backoffLimit := r.BackoffLimit - if job.Spec.BackoffLimit != nil { - backoffLimit = *job.Spec.BackoffLimit - } - - if job.Status.Failed >= backoffLimit { + if isJobConditionTrue(job, batchv1.JobFailed) { logger.Info("migration job failed, will delete and retry", "job", jobName, "version", desiredVersion) // Set condition so kubectl describe shows the failure. @@ -238,6 +233,18 @@ func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) ( return ctrl.Result{RequeueAfter: 10 * time.Second}, nil } +// isJobConditionTrue returns true if the Job has a condition of the given type +// with status True. This is more reliable than comparing status counters because +// the Job controller sets conditions atomically when it makes its final decision. +func isJobConditionTrue(job *batchv1.Job, conditionType batchv1.JobConditionType) bool { + for _, c := range job.Status.Conditions { + if c.Type == conditionType && c.Status == corev1.ConditionTrue { + return true + } + } + return false +} + // isMemoryDatastore checks if the Deployment is using the memory datastore // (no database migration needed). func isMemoryDatastore(container *corev1.Container) bool { diff --git a/operator/internal/controller/migration_controller_test.go b/operator/internal/controller/migration_controller_test.go index be37fe3c..c410aaff 100644 --- a/operator/internal/controller/migration_controller_test.go +++ b/operator/internal/controller/migration_controller_test.go @@ -217,6 +217,12 @@ func TestReconcile_JobSucceeded_UpdatesConfigMapAndScalesUp(t *testing.T) { }, Status: batchv1.JobStatus{ Succeeded: 1, + Conditions: []batchv1.JobCondition{ + { + Type: batchv1.JobComplete, + Status: corev1.ConditionTrue, + }, + }, }, } @@ -296,6 +302,12 @@ func TestReconcile_JobFailed_SetsRetryAnnotationAndRequeues(t *testing.T) { }, Status: batchv1.JobStatus{ Failed: 3, + Conditions: []batchv1.JobCondition{ + { + Type: batchv1.JobFailed, + Status: corev1.ConditionTrue, + }, + }, }, } @@ -522,6 +534,12 @@ func TestReconcile_StaleJob_DeletedAndRequeued(t *testing.T) { }, Status: batchv1.JobStatus{ Succeeded: 1, + Conditions: []batchv1.JobCondition{ + { + Type: batchv1.JobComplete, + Status: corev1.ConditionTrue, + }, + }, }, } From 174f3a1848e7641f72e34397cf4b0f57e508d42a Mon Sep 17 00:00:00 2001 From: Ed Milic <edmilic@gmail.com> Date: Sat, 11 Apr 2026 14:39:10 -0400 Subject: [PATCH 18/70] test: add helm-unittest tests for operator mode 20 new tests across 4 files covering the operator-enabled code paths that previously had no unit test coverage: - deployment: annotations, replicas=0, autoscaling conflict, initContainers gating - job: template not rendered when operator enabled - serviceaccount: migration SA creation, custom names, IRSA annotations - rbac: legacy Role/RoleBinding not rendered when operator enabled --- .../openfga/tests/operator_mode_job_test.yaml | 28 ++++ .../tests/operator_mode_rbac_test.yaml | 25 ++++ .../operator_mode_serviceaccount_test.yaml | 66 +++++++++ charts/openfga/tests/operator_mode_test.yaml | 134 ++++++++++++++++++ 4 files changed, 253 insertions(+) create mode 100644 charts/openfga/tests/operator_mode_job_test.yaml create mode 100644 charts/openfga/tests/operator_mode_rbac_test.yaml create mode 100644 charts/openfga/tests/operator_mode_serviceaccount_test.yaml create mode 100644 charts/openfga/tests/operator_mode_test.yaml diff --git a/charts/openfga/tests/operator_mode_job_test.yaml b/charts/openfga/tests/operator_mode_job_test.yaml new file mode 100644 index 00000000..31d57607 --- /dev/null +++ b/charts/openfga/tests/operator_mode_job_test.yaml @@ -0,0 +1,28 @@ +suite: operator mode - job template +templates: + - templates/job.yaml +tests: + - it: should not render migration job when operator is enabled + set: + operator.enabled: true + migration.enabled: true + datastore.engine: postgres + datastore.uri: "postgres://localhost/openfga" + datastore.applyMigrations: true + datastore.migrationType: job + asserts: + - hasDocuments: + count: 0 + + - it: should render migration job when operator is disabled + set: + operator.enabled: false + datastore.engine: postgres + datastore.uri: "postgres://localhost/openfga" + datastore.applyMigrations: true + datastore.migrationType: job + asserts: + - hasDocuments: + count: 1 + - isKind: + of: Job diff --git a/charts/openfga/tests/operator_mode_rbac_test.yaml b/charts/openfga/tests/operator_mode_rbac_test.yaml new file mode 100644 index 00000000..bb60846c --- /dev/null +++ b/charts/openfga/tests/operator_mode_rbac_test.yaml @@ -0,0 +1,25 @@ +suite: operator mode - RBAC +templates: + - templates/rbac.yaml +tests: + - it: should not render legacy RBAC when operator is enabled + set: + operator.enabled: true + serviceAccount.create: true + asserts: + - hasDocuments: + count: 0 + + - it: should render legacy RBAC when operator is disabled + set: + operator.enabled: false + serviceAccount.create: true + asserts: + - hasDocuments: + count: 2 + - isKind: + of: Role + documentIndex: 0 + - isKind: + of: RoleBinding + documentIndex: 1 diff --git a/charts/openfga/tests/operator_mode_serviceaccount_test.yaml b/charts/openfga/tests/operator_mode_serviceaccount_test.yaml new file mode 100644 index 00000000..cbeab1a0 --- /dev/null +++ b/charts/openfga/tests/operator_mode_serviceaccount_test.yaml @@ -0,0 +1,66 @@ +suite: operator mode - service accounts +templates: + - templates/serviceaccount.yaml +tests: + - it: should render migration service account when operator is enabled + set: + operator.enabled: true + migration.enabled: true + migration.serviceAccount.create: true + serviceAccount.create: true + asserts: + - hasDocuments: + count: 2 + - isKind: + of: ServiceAccount + documentIndex: 1 + - equal: + path: metadata.name + value: RELEASE-NAME-openfga-migration + documentIndex: 1 + + - it: should not render migration service account when operator is disabled + set: + operator.enabled: false + serviceAccount.create: true + asserts: + - hasDocuments: + count: 1 + + - it: should not render migration service account when migration SA creation is disabled + set: + operator.enabled: true + migration.enabled: true + migration.serviceAccount.create: false + migration.serviceAccount.name: external-sa + serviceAccount.create: true + asserts: + - hasDocuments: + count: 1 + + - it: should render migration service account with custom annotations + set: + operator.enabled: true + migration.enabled: true + migration.serviceAccount.create: true + migration.serviceAccount.annotations: + eks.amazonaws.com/role-arn: "arn:aws:iam::123456789012:role/openfga-migrator" + serviceAccount.create: true + asserts: + - equal: + path: metadata.annotations["eks.amazonaws.com/role-arn"] + value: "arn:aws:iam::123456789012:role/openfga-migrator" + documentIndex: 1 + + - it: should use custom migration service account name + set: + operator.enabled: true + migration.enabled: true + migration.serviceAccount.create: true + migration.serviceAccount.name: my-migrator + serviceAccount.create: true + asserts: + - equal: + path: metadata.name + value: my-migrator + documentIndex: 1 diff --git a/charts/openfga/tests/operator_mode_test.yaml b/charts/openfga/tests/operator_mode_test.yaml new file mode 100644 index 00000000..31d06c1b --- /dev/null +++ b/charts/openfga/tests/operator_mode_test.yaml @@ -0,0 +1,134 @@ +suite: operator mode +templates: + - templates/deployment.yaml +tests: + # --- Deployment annotations --- + - it: should set operator annotations when operator and migration are enabled + set: + operator.enabled: true + migration.enabled: true + replicaCount: 3 + datastore.engine: postgres + asserts: + - equal: + path: metadata.annotations["openfga.dev/migration-enabled"] + value: "true" + - equal: + path: metadata.annotations["openfga.dev/desired-replicas"] + value: "3" + - equal: + path: metadata.annotations["openfga.dev/migration-service-account"] + value: RELEASE-NAME-openfga-migration + + - it: should not set operator annotations when operator is disabled + set: + operator.enabled: false + annotations: + custom: value + asserts: + - isNull: + path: metadata.annotations["openfga.dev/migration-enabled"] + - isNull: + path: metadata.annotations["openfga.dev/desired-replicas"] + + - it: should set desired-replicas to 1 for memory datastore + set: + operator.enabled: true + migration.enabled: true + replicaCount: 5 + datastore.engine: memory + asserts: + - equal: + path: metadata.annotations["openfga.dev/desired-replicas"] + value: "1" + + - it: should use custom migration service account name when set + set: + operator.enabled: true + migration.enabled: true + datastore.engine: postgres + migration.serviceAccount.name: my-custom-sa + asserts: + - equal: + path: metadata.annotations["openfga.dev/migration-service-account"] + value: my-custom-sa + + - it: should not set migration-service-account annotation when SA creation is disabled and no name set + set: + operator.enabled: true + migration.enabled: true + datastore.engine: postgres + migration.serviceAccount.create: false + asserts: + - isNull: + path: metadata.annotations["openfga.dev/migration-service-account"] + + # --- Replica count --- + - it: should set replicas to 0 when operator is enabled with database datastore + set: + operator.enabled: true + migration.enabled: true + replicaCount: 3 + datastore.engine: postgres + asserts: + - equal: + path: spec.replicas + value: 0 + + - it: should set replicas to 1 when operator is enabled with memory datastore + set: + operator.enabled: true + migration.enabled: true + replicaCount: 5 + datastore.engine: memory + asserts: + - equal: + path: spec.replicas + value: 1 + + - it: should set replicas to replicaCount when operator is disabled + set: + operator.enabled: false + replicaCount: 5 + datastore.engine: postgres + asserts: + - equal: + path: spec.replicas + value: 5 + + # --- Autoscaling conflict --- + - it: should fail when operator and autoscaling are both enabled + set: + operator.enabled: true + migration.enabled: true + autoscaling.enabled: true + datastore.engine: postgres + asserts: + - failedTemplate: + errorMessage: "operator.enabled and autoscaling.enabled cannot both be true" + + # --- initContainers gating --- + - it: should not render migration initContainers when operator is enabled + set: + operator.enabled: true + migration.enabled: true + datastore.engine: postgres + datastore.uri: "postgres://localhost/openfga" + datastore.applyMigrations: true + datastore.waitForMigrations: true + datastore.migrationType: job + asserts: + - isNull: + path: spec.template.spec.initContainers + + - it: should render migration initContainers when operator is disabled + set: + operator.enabled: false + datastore.engine: postgres + datastore.uri: "postgres://localhost/openfga" + datastore.applyMigrations: true + datastore.waitForMigrations: true + datastore.migrationType: job + asserts: + - isNotNull: + path: spec.template.spec.initContainers From b63aacc40dd2f60ab219e0c5d4149a81b269f6f4 Mon Sep 17 00:00:00 2001 From: Ed Milic <edmilic@gmail.com> Date: Sat, 11 Apr 2026 14:50:27 -0400 Subject: [PATCH 19/70] fix: inherit resource limits in operator migration Job The old Helm-templated migration Job uses datastore.migrations.resources for resource limits, but the operator-built Job had none. Inherit the main container's Resources to maintain parity and prevent unbounded resource consumption during migrations. --- operator/internal/controller/helpers.go | 1 + 1 file changed, 1 insertion(+) diff --git a/operator/internal/controller/helpers.go b/operator/internal/controller/helpers.go index 7a4ac2ab..794210b4 100644 --- a/operator/internal/controller/helpers.go +++ b/operator/internal/controller/helpers.go @@ -154,6 +154,7 @@ func buildMigrationJob( Args: []string{"migrate"}, Env: mainContainer.Env, EnvFrom: mainContainer.EnvFrom, + Resources: mainContainer.Resources, VolumeMounts: mainContainer.VolumeMounts, SecurityContext: mainContainer.SecurityContext, }, From 13c1f17ca91758bd7b4ccb2e11939ad2e49c0206 Mon Sep 17 00:00:00 2001 From: Ed Milic <edmilic@gmail.com> Date: Sat, 11 Apr 2026 14:54:33 -0400 Subject: [PATCH 20/70] test: add missing controller unit tests for edge cases Four new tests covering previously untested code paths: - StaleJob_LabelOnlyFallback: version-mismatch detection when Job has only a label (no annotation), exercising the sanitized-label fallback - JobSucceeded_UpdatesExistingConfigMap: ConfigMap update path when a prior version's ConfigMap already exists - ScaleToZero_NilAnnotationsMap: scaleDeploymentToZero correctly stores desired-replicas when the annotation was not previously set - JobInProgress_Requeues: in-progress Job triggers 10s requeue without scaling up or modifying the Deployment --- .../controller/migration_controller_test.go | 246 ++++++++++++++++++ 1 file changed, 246 insertions(+) diff --git a/operator/internal/controller/migration_controller_test.go b/operator/internal/controller/migration_controller_test.go index c410aaff..b5aa702f 100644 --- a/operator/internal/controller/migration_controller_test.go +++ b/operator/internal/controller/migration_controller_test.go @@ -617,6 +617,252 @@ func TestReconcile_MigrationNotEnabled_Skips(t *testing.T) { } } +func TestReconcile_StaleJob_LabelOnlyFallback_DeletedAndRequeued(t *testing.T) { + // Given: a Deployment at v1.15.0 with an existing Job that only has a label (no annotation). + dep := newTestDeployment("openfga", "default", "openfga/openfga:v1.15.0", 0) + dep.Annotations[AnnotationDesiredReplicas] = "3" + + staleJob := &batchv1.Job{ + ObjectMeta: metav1.ObjectMeta{ + Name: "openfga-migrate", + Namespace: "default", + Labels: map[string]string{ + "app.kubernetes.io/version": "v1.14.0", + }, + // No annotation — forces the label-only fallback path. + OwnerReferences: []metav1.OwnerReference{ + { + APIVersion: "apps/v1", + Kind: "Deployment", + Name: "openfga", + UID: "test-uid-123", + }, + }, + }, + Spec: batchv1.JobSpec{ + BackoffLimit: ptr.To(int32(3)), + Template: corev1.PodTemplateSpec{ + Spec: corev1.PodSpec{ + Containers: []corev1.Container{{Name: "migrate", Image: "openfga/openfga:v1.14.0"}}, + RestartPolicy: corev1.RestartPolicyNever, + }, + }, + }, + } + + r := newReconciler(dep, staleJob) + + // When: reconciling. + result, err := r.Reconcile(context.Background(), ctrl.Request{ + NamespacedName: types.NamespacedName{Name: "openfga", Namespace: "default"}, + }) + + // Then: stale Job should be deleted and requeue requested. + if err != nil { + t.Fatalf("unexpected error: %v", err) + } + if result.RequeueAfter == 0 { + t.Error("expected requeue after deleting stale job") + } + + deletedJob := &batchv1.Job{} + if getErr := r.Get(context.Background(), types.NamespacedName{ + Name: "openfga-migrate", Namespace: "default", + }, deletedJob); getErr == nil { + t.Error("expected stale migration job to be deleted") + } +} + +func TestReconcile_JobSucceeded_UpdatesExistingConfigMap(t *testing.T) { + // Given: a Deployment with a pre-existing ConfigMap from v1.13.0 and a succeeded Job for v1.14.0. + dep := newTestDeployment("openfga", "default", "openfga/openfga:v1.14.0", 0) + dep.Annotations[AnnotationDesiredReplicas] = "3" + + existingCM := &corev1.ConfigMap{ + ObjectMeta: metav1.ObjectMeta{ + Name: "openfga-migration-status", + Namespace: "default", + Labels: map[string]string{ + LabelPartOf: LabelPartOfValue, + LabelComponent: "migration", + "app.kubernetes.io/managed-by": "openfga-operator", + }, + OwnerReferences: []metav1.OwnerReference{ + { + APIVersion: "apps/v1", + Kind: "Deployment", + Name: "openfga", + UID: "test-uid-123", + }, + }, + }, + Data: map[string]string{ + "version": "v1.13.0", + "migratedAt": "2026-04-01T12:00:00Z", + "jobName": "openfga-migrate", + }, + } + + job := &batchv1.Job{ + ObjectMeta: metav1.ObjectMeta{ + Name: "openfga-migrate", + Namespace: "default", + OwnerReferences: []metav1.OwnerReference{ + { + APIVersion: "apps/v1", + Kind: "Deployment", + Name: "openfga", + UID: "test-uid-123", + }, + }, + }, + Spec: batchv1.JobSpec{ + BackoffLimit: ptr.To(int32(3)), + Template: corev1.PodTemplateSpec{ + Spec: corev1.PodSpec{ + Containers: []corev1.Container{{Name: "migrate", Image: "openfga/openfga:v1.14.0"}}, + RestartPolicy: corev1.RestartPolicyNever, + }, + }, + }, + Status: batchv1.JobStatus{ + Succeeded: 1, + Conditions: []batchv1.JobCondition{ + { + Type: batchv1.JobComplete, + Status: corev1.ConditionTrue, + }, + }, + }, + } + + r := newReconciler(dep, existingCM, job) + + // When: reconciling. + _, err := r.Reconcile(context.Background(), ctrl.Request{ + NamespacedName: types.NamespacedName{Name: "openfga", Namespace: "default"}, + }) + + // Then: no error. + if err != nil { + t.Fatalf("unexpected error: %v", err) + } + + // Verify ConfigMap was updated to v1.14.0. + cm := &corev1.ConfigMap{} + if getErr := r.Get(context.Background(), types.NamespacedName{ + Name: "openfga-migration-status", Namespace: "default", + }, cm); getErr != nil { + t.Fatalf("expected ConfigMap to exist: %v", getErr) + } + if cm.Data["version"] != "v1.14.0" { + t.Errorf("expected version v1.14.0 in ConfigMap, got %s", cm.Data["version"]) + } +} + +func TestReconcile_ScaleToZero_NilAnnotationsMap(t *testing.T) { + // Given: a Deployment with nil Annotations map and replicas > 0. + dep := newTestDeployment("openfga", "default", "openfga/openfga:v1.14.0", 3) + dep.Annotations = nil + // Re-add the required annotation via a fresh map — but test that scaleDeploymentToZero + // handles nil gracefully by setting it only via the migration-enabled annotation. + dep.Annotations = map[string]string{ + AnnotationMigrationEnabled: "true", + } + + r := newReconciler(dep) + + // When: reconciling — this will call scaleDeploymentToZero which must handle + // the case where AnnotationDesiredReplicas is not yet set. + result, err := r.Reconcile(context.Background(), ctrl.Request{ + NamespacedName: types.NamespacedName{Name: "openfga", Namespace: "default"}, + }) + + // Then: no error, Job created. + if err != nil { + t.Fatalf("unexpected error: %v", err) + } + if result.RequeueAfter == 0 { + t.Error("expected requeue after creating job") + } + + // Verify Deployment was scaled to 0 and desired-replicas annotation was preserved. + updated := &appsv1.Deployment{} + if getErr := r.Get(context.Background(), types.NamespacedName{ + Name: "openfga", Namespace: "default", + }, updated); getErr != nil { + t.Fatalf("getting deployment: %v", getErr) + } + if *updated.Spec.Replicas != 0 { + t.Errorf("expected 0 replicas, got %d", *updated.Spec.Replicas) + } + if updated.Annotations[AnnotationDesiredReplicas] != "3" { + t.Errorf("expected desired-replicas=3, got %s", updated.Annotations[AnnotationDesiredReplicas]) + } +} + +func TestReconcile_JobInProgress_Requeues(t *testing.T) { + // Given: a Deployment with a running Job (no conditions set yet). + dep := newTestDeployment("openfga", "default", "openfga/openfga:v1.14.0", 0) + dep.Annotations[AnnotationDesiredReplicas] = "3" + + job := &batchv1.Job{ + ObjectMeta: metav1.ObjectMeta{ + Name: "openfga-migrate", + Namespace: "default", + Annotations: map[string]string{ + "openfga.dev/desired-version": "v1.14.0", + }, + OwnerReferences: []metav1.OwnerReference{ + { + APIVersion: "apps/v1", + Kind: "Deployment", + Name: "openfga", + UID: "test-uid-123", + }, + }, + }, + Spec: batchv1.JobSpec{ + BackoffLimit: ptr.To(int32(3)), + Template: corev1.PodTemplateSpec{ + Spec: corev1.PodSpec{ + Containers: []corev1.Container{{Name: "migrate", Image: "openfga/openfga:v1.14.0"}}, + RestartPolicy: corev1.RestartPolicyNever, + }, + }, + }, + Status: batchv1.JobStatus{ + Active: 1, + }, + } + + r := newReconciler(dep, job) + + // When: reconciling. + result, err := r.Reconcile(context.Background(), ctrl.Request{ + NamespacedName: types.NamespacedName{Name: "openfga", Namespace: "default"}, + }) + + // Then: no error, requeue after 10s to poll progress. + if err != nil { + t.Fatalf("unexpected error: %v", err) + } + if result.RequeueAfter != 10*time.Second { + t.Errorf("expected 10s requeue for in-progress job, got %v", result.RequeueAfter) + } + + // Verify Deployment still at 0 replicas. + updated := &appsv1.Deployment{} + if getErr := r.Get(context.Background(), types.NamespacedName{ + Name: "openfga", Namespace: "default", + }, updated); getErr != nil { + t.Fatalf("getting deployment: %v", getErr) + } + if *updated.Spec.Replicas != 0 { + t.Errorf("expected 0 replicas while job in progress, got %d", *updated.Spec.Replicas) + } +} + func TestExtractImageTag(t *testing.T) { tests := []struct { image string From 2b9de0b49c1cc820341dadd223ef552ccdb6c014 Mon Sep 17 00:00:00 2001 From: Ed Milic <edmilic@gmail.com> Date: Sat, 11 Apr 2026 15:07:33 -0400 Subject: [PATCH 21/70] fix: wire migration Job flags (backoff/deadline/TTL) through Helm values The README documented these flags as configurable via the subchart values.yaml, but only --leader-elect and --watch-namespace were actually wired. Add migrationJob.backoffLimit, activeDeadlineSeconds, and ttlSecondsAfterFinished values with matching args in the operator Deployment template. --- charts/openfga-operator/templates/deployment.yaml | 3 +++ charts/openfga-operator/values.yaml | 8 ++++++++ 2 files changed, 11 insertions(+) diff --git a/charts/openfga-operator/templates/deployment.yaml b/charts/openfga-operator/templates/deployment.yaml index ecad0916..4ef3b804 100644 --- a/charts/openfga-operator/templates/deployment.yaml +++ b/charts/openfga-operator/templates/deployment.yaml @@ -43,6 +43,9 @@ spec: {{- if .Values.watchNamespace }} - --watch-namespace={{ .Values.watchNamespace }} {{- end }} + - --backoff-limit={{ .Values.migrationJob.backoffLimit }} + - --active-deadline-seconds={{ .Values.migrationJob.activeDeadlineSeconds }} + - --ttl-seconds-after-finished={{ .Values.migrationJob.ttlSecondsAfterFinished }} env: - name: POD_NAMESPACE valueFrom: diff --git a/charts/openfga-operator/values.yaml b/charts/openfga-operator/values.yaml index 4b308a21..56a0b5b9 100644 --- a/charts/openfga-operator/values.yaml +++ b/charts/openfga-operator/values.yaml @@ -42,6 +42,14 @@ leaderElection: # -- Enable leader election for controller manager. enabled: true +migrationJob: + # -- Number of pod failures before a migration Job is considered failed. + backoffLimit: 3 + # -- Maximum wall-clock seconds a migration Job can run before being terminated. + activeDeadlineSeconds: 300 + # -- Seconds to keep completed/failed Job pods for log inspection before garbage collection. + ttlSecondsAfterFinished: 300 + resources: {} # requests: # cpu: 10m From f4048d3e77b0c5c28929e858404b495a9ead3730 Mon Sep 17 00:00:00 2001 From: Ed Milic <edmilic@gmail.com> Date: Sat, 11 Apr 2026 15:08:55 -0400 Subject: [PATCH 22/70] fix: document namespaceOverride in operator subchart values.yaml The _helpers.tpl namespace template already supported namespaceOverride but the value was undeclared in values.yaml, making it undiscoverable. --- charts/openfga-operator/values.yaml | 3 +++ 1 file changed, 3 insertions(+) diff --git a/charts/openfga-operator/values.yaml b/charts/openfga-operator/values.yaml index 56a0b5b9..6a4c1659 100644 --- a/charts/openfga-operator/values.yaml +++ b/charts/openfga-operator/values.yaml @@ -9,6 +9,9 @@ image: imagePullSecrets: [] nameOverride: "" fullnameOverride: "" +# -- Override the namespace for all operator resources. +# Useful when the parent chart deploys subcharts into a different namespace. +namespaceOverride: "" serviceAccount: # -- Specifies whether a service account should be created. From a80d4bfc1da000bd38e0322e70b5a0c5ef8960a8 Mon Sep 17 00:00:00 2001 From: Ed Milic <edmilic@gmail.com> Date: Sat, 11 Apr 2026 15:10:18 -0400 Subject: [PATCH 23/70] fix: rename misleading ScaleToZero test to match actual behavior The test validates that scaleDeploymentToZero stores the current replica count in the desired-replicas annotation before zeroing, not that it handles a nil annotations map. --- operator/internal/controller/migration_controller_test.go | 8 +++----- 1 file changed, 3 insertions(+), 5 deletions(-) diff --git a/operator/internal/controller/migration_controller_test.go b/operator/internal/controller/migration_controller_test.go index b5aa702f..a2edbf68 100644 --- a/operator/internal/controller/migration_controller_test.go +++ b/operator/internal/controller/migration_controller_test.go @@ -760,12 +760,10 @@ func TestReconcile_JobSucceeded_UpdatesExistingConfigMap(t *testing.T) { } } -func TestReconcile_ScaleToZero_NilAnnotationsMap(t *testing.T) { - // Given: a Deployment with nil Annotations map and replicas > 0. +func TestReconcile_ScaleToZero_StoresDesiredReplicas(t *testing.T) { + // Given: a Deployment with replicas > 0 and no desired-replicas annotation yet. + // scaleDeploymentToZero should store the current replica count before zeroing. dep := newTestDeployment("openfga", "default", "openfga/openfga:v1.14.0", 3) - dep.Annotations = nil - // Re-add the required annotation via a fresh map — but test that scaleDeploymentToZero - // handles nil gracefully by setting it only via the migration-enabled annotation. dep.Annotations = map[string]string{ AnnotationMigrationEnabled: "true", } From a9411bc997412658b4fac72e33121dc8d65165cd Mon Sep 17 00:00:00 2001 From: Ed Milic <edmilic@gmail.com> Date: Sat, 11 Apr 2026 15:14:29 -0400 Subject: [PATCH 24/70] fix: require explicit serviceAccount.name when create=false Falling back to the "default" ServiceAccount would silently grant operator RBAC permissions (Deployment patch, Job create/delete) to a shared SA. Require an explicit name so the user makes a deliberate choice. --- charts/openfga-operator/templates/_helpers.tpl | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/charts/openfga-operator/templates/_helpers.tpl b/charts/openfga-operator/templates/_helpers.tpl index 70d6e4c4..f63057d6 100644 --- a/charts/openfga-operator/templates/_helpers.tpl +++ b/charts/openfga-operator/templates/_helpers.tpl @@ -67,6 +67,6 @@ Create the name of the service account to use {{- if .Values.serviceAccount.create }} {{- default (include "openfga-operator.fullname" .) .Values.serviceAccount.name }} {{- else }} -{{- default "default" .Values.serviceAccount.name }} +{{- required "serviceAccount.name must be set when serviceAccount.create=false" .Values.serviceAccount.name }} {{- end }} {{- end }} From bd2725a45753528d0ccf427f3ee95824707f0f5a Mon Sep 17 00:00:00 2001 From: Ed Milic <edmilic@gmail.com> Date: Sat, 11 Apr 2026 15:15:29 -0400 Subject: [PATCH 25/70] docs: update chart structure --- docs/adr/004-operator-deployment-model.md | 6 +++--- 1 file changed, 3 insertions(+), 3 deletions(-) diff --git a/docs/adr/004-operator-deployment-model.md b/docs/adr/004-operator-deployment-model.md index 603bf3a2..a5a693c6 100644 --- a/docs/adr/004-operator-deployment-model.md +++ b/docs/adr/004-operator-deployment-model.md @@ -66,7 +66,7 @@ The operator will be distributed as a **conditional Helm subchart dependency** o ### Chart Structure -``` +```text helm-charts/ ├── charts/ │ ├── openfga/ # Main chart (existing) @@ -81,8 +81,8 @@ helm-charts/ │ ├── templates/ │ │ ├── deployment.yaml │ │ ├── serviceaccount.yaml -│ │ ├── clusterrole.yaml -│ │ └── clusterrolebinding.yaml +│ │ ├── role.yaml +│ │ └── rolebinding.yaml │ └── crds/ # CRDs added in Stages 2-4 │ ├── fgastore.yaml │ ├── fgamodel.yaml From 962bb3f51107bee7c61bd0e8df446ef507a79849 Mon Sep 17 00:00:00 2001 From: Ed Milic <edmilic@gmail.com> Date: Sat, 11 Apr 2026 15:30:09 -0400 Subject: [PATCH 26/70] fix: handle AlreadyExists on migration Job creation gracefully When concurrent reconciles race between the GET and CREATE, the second create returns AlreadyExists. Treat this as benign and requeue to poll the existing Job instead of returning a hard error that produces noisy reconcile failures in the controller logs. --- operator/internal/controller/migration_controller.go | 5 +++++ 1 file changed, 5 insertions(+) diff --git a/operator/internal/controller/migration_controller.go b/operator/internal/controller/migration_controller.go index 75cbf199..de044d6b 100644 --- a/operator/internal/controller/migration_controller.go +++ b/operator/internal/controller/migration_controller.go @@ -137,6 +137,11 @@ func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) ( } } if createErr := r.Create(ctx, job); createErr != nil { + if apierrors.IsAlreadyExists(createErr) { + // A concurrent reconcile already created the Job; requeue to pick it up. + logger.V(1).Info("migration job already exists, will recheck", "job", jobName) + return ctrl.Result{RequeueAfter: 5 * time.Second}, nil + } return ctrl.Result{}, fmt.Errorf("creating migration job: %w", createErr) } logger.Info("created migration job", "job", jobName, "version", desiredVersion) From 81e40ba753a70cada555d1284b4320d9f2db3cc8 Mon Sep 17 00:00:00 2001 From: Ed Milic <edmilic@gmail.com> Date: Sat, 11 Apr 2026 15:57:37 -0400 Subject: [PATCH 27/70] fix: harden openfga-operator chart security and quality defaults Address review findings: add allowPrivilegeEscalation: false for restricted PSS compliance, set default resource requests/limits, add values.schema.json validation, use stable selectorLabels on pod template to prevent spurious rollouts, add .helmignore, and improve Chart.yaml metadata and NOTES.txt with migration commands. --- charts/openfga-operator/.helmignore | 18 ++++ charts/openfga-operator/Chart.yaml | 6 ++ charts/openfga-operator/templates/NOTES.txt | 6 ++ .../templates/deployment.yaml | 2 +- charts/openfga-operator/values.schema.json | 100 ++++++++++++++++++ charts/openfga-operator/values.yaml | 13 +-- .../controller/migration_controller.go | 8 +- 7 files changed, 144 insertions(+), 9 deletions(-) create mode 100644 charts/openfga-operator/.helmignore create mode 100644 charts/openfga-operator/values.schema.json diff --git a/charts/openfga-operator/.helmignore b/charts/openfga-operator/.helmignore new file mode 100644 index 00000000..edf9e7ef --- /dev/null +++ b/charts/openfga-operator/.helmignore @@ -0,0 +1,18 @@ +# Patterns to ignore when building packages. +.DS_Store +.git +.gitignore +.bzr +.bzrignore +.hg +.hgignore +.svn +*.swp +*.bak +*.tmp +*.orig +*~ +.project +.idea +*.tmproj +.vscode diff --git a/charts/openfga-operator/Chart.yaml b/charts/openfga-operator/Chart.yaml index 1bdacb03..95da06ba 100644 --- a/charts/openfga-operator/Chart.yaml +++ b/charts/openfga-operator/Chart.yaml @@ -9,5 +9,11 @@ appVersion: "0.1.0" home: "https://openfga.github.io/helm-charts" icon: https://github.com/openfga/community/raw/main/brand-assets/icon/color/openfga-icon-color.svg +maintainers: + - name: OpenFGA Authors + url: https://github.com/openfga +sources: + - https://github.com/openfga/helm-charts + annotations: artifacthub.io/license: Apache-2.0 diff --git a/charts/openfga-operator/templates/NOTES.txt b/charts/openfga-operator/templates/NOTES.txt index c2e09f9a..bcfcf801 100644 --- a/charts/openfga-operator/templates/NOTES.txt +++ b/charts/openfga-operator/templates/NOTES.txt @@ -8,3 +8,9 @@ To check operator status: To view operator logs: kubectl logs --namespace {{ include "openfga-operator.namespace" . }} -l "app.kubernetes.io/name={{ include "openfga-operator.name" . }}" + +To check migration status: + kubectl get configmap -n {{ include "openfga-operator.namespace" . }} -l app.kubernetes.io/managed-by=openfga-operator + +To inspect migration jobs: + kubectl get jobs -n {{ include "openfga-operator.namespace" . }} -l app.kubernetes.io/part-of=openfga,app.kubernetes.io/component=migration diff --git a/charts/openfga-operator/templates/deployment.yaml b/charts/openfga-operator/templates/deployment.yaml index 4ef3b804..5b83a6bd 100644 --- a/charts/openfga-operator/templates/deployment.yaml +++ b/charts/openfga-operator/templates/deployment.yaml @@ -17,7 +17,7 @@ spec: {{- toYaml . | nindent 8 }} {{- end }} labels: - {{- include "openfga-operator.labels" . | nindent 8 }} + {{- include "openfga-operator.selectorLabels" . | nindent 8 }} spec: {{- with .Values.imagePullSecrets }} imagePullSecrets: diff --git a/charts/openfga-operator/values.schema.json b/charts/openfga-operator/values.schema.json new file mode 100644 index 00000000..470b265a --- /dev/null +++ b/charts/openfga-operator/values.schema.json @@ -0,0 +1,100 @@ +{ + "$schema": "https://json-schema.org/draft/2020-12/schema", + "type": "object", + "properties": { + "replicaCount": { + "type": "integer", + "minimum": 1 + }, + "image": { + "type": "object", + "properties": { + "repository": { + "type": "string", + "minLength": 1 + }, + "pullPolicy": { + "type": "string", + "enum": ["Always", "IfNotPresent", "Never"] + }, + "tag": { + "type": "string" + } + }, + "required": ["repository"] + }, + "imagePullSecrets": { + "type": "array", + "items": { + "type": "object", + "properties": { + "name": { "type": "string" } + }, + "required": ["name"] + } + }, + "nameOverride": { "type": "string" }, + "fullnameOverride": { "type": "string" }, + "namespaceOverride": { "type": "string" }, + "serviceAccount": { + "type": "object", + "properties": { + "create": { "type": "boolean" }, + "annotations": { "type": "object" }, + "name": { "type": "string" } + } + }, + "podAnnotations": { "type": "object" }, + "podSecurityContext": { "type": "object" }, + "securityContext": { "type": "object" }, + "watchNamespace": { "type": "string" }, + "leaderElection": { + "type": "object", + "properties": { + "enabled": { "type": "boolean" } + } + }, + "migrationJob": { + "type": "object", + "properties": { + "backoffLimit": { + "type": "integer", + "minimum": 0 + }, + "activeDeadlineSeconds": { + "type": "integer", + "minimum": 1 + }, + "ttlSecondsAfterFinished": { + "type": "integer", + "minimum": 0 + } + } + }, + "resources": { "type": "object" }, + "podDisruptionBudget": { + "type": "object", + "properties": { + "enabled": { "type": "boolean" }, + "minAvailable": { + "oneOf": [ + { "type": "string" }, + { "type": "integer", "minimum": 0 } + ] + }, + "maxUnavailable": { + "oneOf": [ + { "type": "string" }, + { "type": "integer", "minimum": 0 } + ] + } + } + }, + "nodeSelector": { "type": "object" }, + "tolerations": { + "type": "array", + "items": { "type": "object" } + }, + "affinity": { "type": "object" } + } +} diff --git a/charts/openfga-operator/values.yaml b/charts/openfga-operator/values.yaml index 6a4c1659..db2747b4 100644 --- a/charts/openfga-operator/values.yaml +++ b/charts/openfga-operator/values.yaml @@ -30,6 +30,7 @@ podSecurityContext: type: RuntimeDefault securityContext: + allowPrivilegeEscalation: false capabilities: drop: - ALL @@ -53,12 +54,12 @@ migrationJob: # -- Seconds to keep completed/failed Job pods for log inspection before garbage collection. ttlSecondsAfterFinished: 300 -resources: {} - # requests: - # cpu: 10m - # memory: 64Mi - # limits: - # memory: 128Mi +resources: + requests: + cpu: 10m + memory: 64Mi + limits: + memory: 128Mi podDisruptionBudget: # -- Enable a PodDisruptionBudget for the operator. diff --git a/operator/internal/controller/migration_controller.go b/operator/internal/controller/migration_controller.go index de044d6b..2fea3b4f 100644 --- a/operator/internal/controller/migration_controller.go +++ b/operator/internal/controller/migration_controller.go @@ -48,7 +48,7 @@ func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) ( } // 2. Skip if migration is not opted-in via annotation. - if deployment.Annotations[AnnotationMigrationEnabled] != "true" { + if len(deployment.Annotations) == 0 || deployment.Annotations[AnnotationMigrationEnabled] != "true" { logger.V(1).Info("migration not enabled for this deployment, skipping") return ctrl.Result{}, nil } @@ -61,7 +61,7 @@ func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) ( } desiredVersion := extractImageTag(mainContainer.Image) - // 3. Skip migration for memory datastore — just ensure the Deployment is scaled up. + // 3b. Skip migration for memory datastore — just ensure the Deployment is scaled up. if isMemoryDatastore(mainContainer) { logger.V(1).Info("memory datastore detected, skipping migration") if _, scaleErr := ensureDeploymentScaled(ctx, r.Client, deployment); scaleErr != nil { @@ -252,6 +252,10 @@ func isJobConditionTrue(job *batchv1.Job, conditionType batchv1.JobConditionType // isMemoryDatastore checks if the Deployment is using the memory datastore // (no database migration needed). +// +// NOTE: This only inspects explicit env vars on the container spec. If +// OPENFGA_DATASTORE_ENGINE is injected via envFrom (ConfigMap/Secret), it +// will not be detected here and the operator will attempt a migration. func isMemoryDatastore(container *corev1.Container) bool { for _, env := range container.Env { if env.Name == "OPENFGA_DATASTORE_ENGINE" { From e4f18c5c9c9120c7015a9623673d58c135a0d009 Mon Sep 17 00:00:00 2001 From: Ed Milic <edmilic@gmail.com> Date: Sat, 11 Apr 2026 16:56:43 -0400 Subject: [PATCH 28/70] fix: replace scale-to-zero with lookup-based zero-downtime upgrades MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The operator was scaling Deployments to 0 replicas during every migration, causing a full outage on every helm upgrade — a regression from the existing rolling update behavior. OpenFGA already gates readiness on schema version (MinimumSupportedDatastoreSchemaRevision in sqlcommon.IsReady), so new pods naturally block until migration completes while old pods keep serving. Use Helm's lookup function to preserve the live replica count on upgrade (falling back to replicas: 0 on fresh install where no Deployment exists). Remove scaleDeploymentToZero from the operator reconcile loop. Update ADR-002 to document the rationale and the readiness gate dependency. --- charts/openfga/templates/deployment.yaml | 20 +++++ charts/openfga/tests/operator_mode_test.yaml | 5 +- docs/adr/002-operator-managed-migrations.md | 77 +++++++++++++------ operator/internal/controller/helpers.go | 30 -------- .../controller/migration_controller.go | 11 +-- .../controller/migration_controller_test.go | 22 +++--- 6 files changed, 89 insertions(+), 76 deletions(-) diff --git a/charts/openfga/templates/deployment.yaml b/charts/openfga/templates/deployment.yaml index a70c93dc..a0da306f 100644 --- a/charts/openfga/templates/deployment.yaml +++ b/charts/openfga/templates/deployment.yaml @@ -23,7 +23,27 @@ spec: {{- if .Values.autoscaling.enabled }} {{- fail "operator.enabled and autoscaling.enabled cannot both be true" }} {{- end }} + {{- /* On upgrade, preserve live replica count so existing pods keep serving (zero-downtime). + On fresh install, lookup returns empty — fall back to 0 so no pods start before migration. + See ADR-002 for rationale. */ -}} + {{- $existing := (lookup "apps/v1" "Deployment" (include "openfga.namespace" .) (include "openfga.fullname" .)) }} + {{- $desiredImage := printf "%s:%s" .Values.image.repository (.Values.image.tag | default .Chart.AppVersion) }} + {{- $currentImage := "" }} + {{- range (($existing).spec).template.spec.containers }} + {{- if eq .name "openfga" }} + {{- $currentImage = .image }} + {{- end }} + {{- end }} + {{- if and $existing (eq $currentImage $desiredImage) }} + replicas: {{ $existing.spec.replicas }} + {{- else if $existing }} + {{- /* Image changed — preserve replicas; OpenFGA's built-in schema version check + (MinimumSupportedDatastoreSchemaRevision) causes readiness to fail until + migration completes, so old pods keep serving while new pods wait. */ -}} + replicas: {{ $existing.spec.replicas }} + {{- else }} replicas: {{ ternary 1 0 (eq .Values.datastore.engine "memory") }} + {{- end }} {{- else if not .Values.autoscaling.enabled }} replicas: {{ ternary 1 .Values.replicaCount (eq .Values.datastore.engine "memory")}} {{- end }} diff --git a/charts/openfga/tests/operator_mode_test.yaml b/charts/openfga/tests/operator_mode_test.yaml index 31d06c1b..2e509151 100644 --- a/charts/openfga/tests/operator_mode_test.yaml +++ b/charts/openfga/tests/operator_mode_test.yaml @@ -64,7 +64,10 @@ tests: path: metadata.annotations["openfga.dev/migration-service-account"] # --- Replica count --- - - it: should set replicas to 0 when operator is enabled with database datastore + # When no live cluster is available (helm template / test), lookup returns empty, + # so the template falls back to replicas: 0 (fresh install behavior). + # On a real cluster, lookup preserves the existing replica count for zero-downtime upgrades. + - it: should set replicas to 0 on fresh install when operator is enabled with database datastore set: operator.enabled: true migration.enabled: true diff --git a/docs/adr/002-operator-managed-migrations.md b/docs/adr/002-operator-managed-migrations.md index 9b86d9d7..ad537992 100644 --- a/docs/adr/002-operator-managed-migrations.md +++ b/docs/adr/002-operator-managed-migrations.md @@ -74,30 +74,37 @@ Replace the Helm hook migration Job and `k8s-wait-for` init container with **ope The operator runs a **migration controller** that reconciles the OpenFGA Deployment: ``` -┌────────────────────────────────────────────────────────┐ -│ Operator Reconciliation │ -│ │ -│ 1. Read Deployment → extract image tag (e.g. v1.14.0) │ -│ 2. Read ConfigMap/openfga-migration-status │ -│ └── "Last migrated version: v1.13.0" │ -│ 3. Versions differ → migration needed │ -│ 4. Create Job/openfga-migrate │ -│ ├── ServiceAccount: openfga-migrator (DDL perms) │ -│ ├── Image: openfga/openfga:v1.14.0 │ -│ ├── Args: ["migrate"] │ -│ └── ttlSecondsAfterFinished: 300 │ -│ 5. Watch Job until succeeded │ -│ 6. Update ConfigMap → "version: v1.14.0" │ -│ 7. Scale Deployment replicas: 0 → 3 │ -│ 8. OpenFGA pods start, serve requests │ -└────────────────────────────────────────────────────────┘ +┌──────────────────────────────────────────────────────────┐ +│ Operator Reconciliation │ +│ │ +│ 1. Read Deployment → extract image tag (e.g. v1.14.0) │ +│ 2. Read ConfigMap/openfga-migration-status │ +│ └── "Last migrated version: v1.13.0" │ +│ 3. Versions differ → migration needed │ +│ 4. Create Job/openfga-migrate │ +│ ├── ServiceAccount: openfga-migrator (DDL perms) │ +│ ├── Image: openfga/openfga:v1.14.0 │ +│ ├── Args: ["migrate"] │ +│ └── ttlSecondsAfterFinished: 300 │ +│ 5. Watch Job until succeeded │ +│ 6. Update ConfigMap → "version: v1.14.0" │ +│ 7. Ensure Deployment at desired replicas │ +│ (fresh install: 0 → N; upgrade: already running) │ +│ 8. New pods pass readiness, serve requests │ +└──────────────────────────────────────────────────────────┘ ``` **Key design decisions within this approach:** -#### Deployment starts at replicas: 0 +#### Zero-downtime upgrades via lookup and readiness gating -The Helm chart renders the Deployment with `replicas: 0` when `operator.enabled: true`. The operator scales it up only after migration succeeds. This is simpler than readiness gates or admission webhooks, and ensures no pods run against an unmigrated schema. +On **fresh install**, the Helm chart renders the Deployment with `replicas: 0` (no existing Deployment found via `lookup`). The operator runs the migration Job and scales the Deployment to the desired replica count afterward. + +On **upgrade**, the chart uses Helm's `lookup` function to read the current replica count from the live Deployment and preserves it. Kubernetes starts a rolling update with the new image. OpenFGA has a **built-in schema version gate**: on startup, each instance calls `IsReady()` which checks the database schema revision against `MinimumSupportedDatastoreSchemaRevision` (via goose). If the schema is behind, the gRPC health endpoint returns `NOT_SERVING`, the readiness probe fails, and Kubernetes does not route traffic to the pod. Old pods continue serving on the migrated schema (OpenFGA migrations are additive/backward-compatible — this is how the existing Helm hook flow has operated for years with rolling updates). Once the operator's migration Job completes, new pods pass readiness and the rolling update proceeds. + +This matches the existing zero-downtime behavior of the non-operator chart. The previous approach (always starting at `replicas: 0`) introduced a full outage on every `helm upgrade` — even for config-only changes — which was a regression from the existing rolling update model. + +**`lookup` caveat:** `helm template` and `--dry-run=client` cannot query the cluster, so `lookup` returns empty and the template falls back to `replicas: 0`. This is correct for CI rendering (no live cluster) and does not affect real installs/upgrades. `--dry-run=server` works correctly. #### Version tracking via ConfigMap @@ -142,13 +149,13 @@ helm install Problems: ArgoCD skips step 4. FluxCD deletes Job in step 4. `--wait` deadlocks between steps 2 and 4. -**After (operator-managed):** +**After (operator-managed, fresh install):** ``` helm install ├── Create ServiceAccount (runtime), ServiceAccount (migrator) ├── Create Secret, Service - ├── Create Deployment (replicas: 0, no init containers) + ├── Create Deployment (replicas: 0 via lookup fallback, no init containers) ├── Create Operator Deployment └── [Helm is done — all resources are regular, no hooks] @@ -159,10 +166,30 @@ Operator starts: │ └── Uses openfga-migrator ServiceAccount │ └── Runs openfga migrate → succeeds ├── Creates ConfigMap with migrated version - └── Scales Deployment to 3 replicas → pods start + └── Scales Deployment 0 → 3 replicas → pods start +``` + +**After (operator-managed, upgrade with new image):** + +``` +helm upgrade + ├── lookup finds existing Deployment at 3 replicas → preserves replicas: 3 + ├── Patches Deployment with new image tag + ├── Kubernetes starts rolling update + │ ├── New pods (v1.14) start → schema is behind → + │ │ readiness fails (gRPC NOT_SERVING) → no traffic routed + │ └── Old pods (v1.13) continue serving traffic + └── [Helm is done] + +Operator reconciles: + ├── Detects image version differs from ConfigMap + ├── Creates Job/openfga-migrate → runs migration + ├── Updates ConfigMap → "version: v1.14.0" + └── New pods pass readiness → rolling update completes + (operator does NOT scale to zero — zero downtime) ``` -No hooks. No init containers. No `k8s-wait-for`. All resources are regular Kubernetes objects. +No hooks. No init containers. No `k8s-wait-for`. No downtime on upgrade. All resources are regular Kubernetes objects. ### What Changes in the Helm Chart @@ -206,10 +233,10 @@ When `operator.enabled: false`, the chart falls back to the current behavior — ### Negative - **Operator is a new runtime dependency** — if the operator pod is unavailable, migrations don't run (but existing running pods are unaffected) -- **Replica scaling model** — starting at `replicas: 0` means a brief period where the Deployment exists but has no pods; monitoring tools may flag this +- **`lookup` limitation** — `helm template` and `--dry-run=client` cannot query the cluster; the template falls back to `replicas: 0` in these contexts. This does not affect real installs/upgrades. - **Two upgrade paths to document** — `operator.enabled: true` (new) vs `operator.enabled: false` (legacy) ### Risks -- **Zero-downtime upgrades** — the initial implementation scales to 0 during migration, causing brief downtime. A future enhancement can support rolling upgrades where the new schema is backward-compatible, but this is explicitly out of scope for Stage 1. +- **Readiness gate relies on OpenFGA's built-in schema check** — the zero-downtime upgrade model depends on `MinimumSupportedDatastoreSchemaRevision` in `pkg/storage/sqlcommon/sqlcommon.go` causing `NOT_SERVING` when the schema is behind. If a future OpenFGA release removes or weakens this check, new pods could serve traffic against an unmigrated schema. This coupling should be documented and monitored across OpenFGA releases. - **ConfigMap as state store** — if the ConfigMap is accidentally deleted, the operator re-runs migration (which is safe — `openfga migrate` is idempotent). This is a feature, not a bug, but should be documented. diff --git a/operator/internal/controller/helpers.go b/operator/internal/controller/helpers.go index 794210b4..7eced335 100644 --- a/operator/internal/controller/helpers.go +++ b/operator/internal/controller/helpers.go @@ -264,33 +264,3 @@ func ensureDeploymentScaled(ctx context.Context, c client.Client, deployment *ap return false, nil } -// scaleDeploymentToZero scales the Deployment to 0 replicas, storing the current -// desired count in an annotation so it can be restored later. -func scaleDeploymentToZero(ctx context.Context, c client.Client, deployment *appsv1.Deployment) error { - if deployment.Spec.Replicas != nil && *deployment.Spec.Replicas == 0 { - return nil // Already at zero. - } - - patch := client.MergeFrom(deployment.DeepCopy()) - - // Store the current desired replica count before zeroing. - currentReplicas := int32(1) - if deployment.Spec.Replicas != nil { - currentReplicas = *deployment.Spec.Replicas - } - - // Only store if not already stored (avoid overwriting with 0 on re-reconciliation). - if _, ok := deployment.Annotations[AnnotationDesiredReplicas]; !ok { - if deployment.Annotations == nil { - deployment.Annotations = make(map[string]string) - } - deployment.Annotations[AnnotationDesiredReplicas] = strconv.FormatInt(int64(currentReplicas), 10) - } - - deployment.Spec.Replicas = ptr.To(int32(0)) - - if err := c.Patch(ctx, deployment, patch); err != nil { - return fmt.Errorf("scaling deployment to 0: %w", err) - } - return nil -} diff --git a/operator/internal/controller/migration_controller.go b/operator/internal/controller/migration_controller.go index 2fea3b4f..6152df23 100644 --- a/operator/internal/controller/migration_controller.go +++ b/operator/internal/controller/migration_controller.go @@ -98,12 +98,7 @@ func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) ( logger.Info("migration needed", "currentVersion", currentVersion, "desiredVersion", desiredVersion) - // 6. Ensure the Deployment is scaled to zero before migrating. - if err := scaleDeploymentToZero(ctx, r.Client, deployment); err != nil { - return ctrl.Result{}, err - } - - // 7. Check retry-after annotation to honor backoff cooldown. + // 6. Check retry-after annotation to honor backoff cooldown. if retryAfter, ok := deployment.Annotations[AnnotationRetryAfter]; ok { retryTime, parseErr := time.Parse(time.RFC3339, retryAfter) if parseErr == nil && time.Now().Before(retryTime) { @@ -113,7 +108,7 @@ func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) ( } } - // 8. Check if a migration Job already exists. + // 7. Check if a migration Job already exists. jobName := migrationJobName(req.Name) job := &batchv1.Job{} err = r.Get(ctx, types.NamespacedName{Name: jobName, Namespace: req.Namespace}, job) @@ -150,7 +145,7 @@ func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) ( return ctrl.Result{}, fmt.Errorf("getting migration job: %w", err) } - // 8b. If the existing Job is for a different version, delete it and recreate. + // 8. If the existing Job is for a different version, delete it and recreate. // Check annotation first (supports digests > 63 chars), fall back to label. jobVersion := job.Annotations["openfga.dev/desired-version"] versionMatch := jobVersion == desiredVersion diff --git a/operator/internal/controller/migration_controller_test.go b/operator/internal/controller/migration_controller_test.go index a2edbf68..a4d62a22 100644 --- a/operator/internal/controller/migration_controller_test.go +++ b/operator/internal/controller/migration_controller_test.go @@ -326,7 +326,7 @@ func TestReconcile_JobFailed_SetsRetryAnnotationAndRequeues(t *testing.T) { t.Errorf("expected 60s requeue, got %v", result.RequeueAfter) } - // Verify Deployment was NOT scaled up — still at 0. + // Verify Deployment replicas unchanged (still at 0 from fresh install). updated := &appsv1.Deployment{} if getErr := r.Get(context.Background(), types.NamespacedName{ Name: "openfga", Namespace: "default", @@ -760,18 +760,19 @@ func TestReconcile_JobSucceeded_UpdatesExistingConfigMap(t *testing.T) { } } -func TestReconcile_ScaleToZero_StoresDesiredReplicas(t *testing.T) { - // Given: a Deployment with replicas > 0 and no desired-replicas annotation yet. - // scaleDeploymentToZero should store the current replica count before zeroing. +func TestReconcile_MigrationNeeded_DoesNotScaleToZero(t *testing.T) { + // Given: a Deployment with replicas > 0 and no migration-status ConfigMap. + // The operator should create the migration Job WITHOUT scaling to zero, + // relying on OpenFGA's built-in schema version check to gate readiness. dep := newTestDeployment("openfga", "default", "openfga/openfga:v1.14.0", 3) dep.Annotations = map[string]string{ AnnotationMigrationEnabled: "true", + AnnotationDesiredReplicas: "3", } r := newReconciler(dep) - // When: reconciling — this will call scaleDeploymentToZero which must handle - // the case where AnnotationDesiredReplicas is not yet set. + // When: reconciling. result, err := r.Reconcile(context.Background(), ctrl.Request{ NamespacedName: types.NamespacedName{Name: "openfga", Namespace: "default"}, }) @@ -784,18 +785,15 @@ func TestReconcile_ScaleToZero_StoresDesiredReplicas(t *testing.T) { t.Error("expected requeue after creating job") } - // Verify Deployment was scaled to 0 and desired-replicas annotation was preserved. + // Verify Deployment replicas were NOT changed — pods keep running during migration. updated := &appsv1.Deployment{} if getErr := r.Get(context.Background(), types.NamespacedName{ Name: "openfga", Namespace: "default", }, updated); getErr != nil { t.Fatalf("getting deployment: %v", getErr) } - if *updated.Spec.Replicas != 0 { - t.Errorf("expected 0 replicas, got %d", *updated.Spec.Replicas) - } - if updated.Annotations[AnnotationDesiredReplicas] != "3" { - t.Errorf("expected desired-replicas=3, got %s", updated.Annotations[AnnotationDesiredReplicas]) + if *updated.Spec.Replicas != 3 { + t.Errorf("expected replicas to remain at 3, got %d", *updated.Spec.Replicas) } } From da8cc981b97191460191c7cb99af254f80e4ab43 Mon Sep 17 00:00:00 2001 From: Ed Milic <edmilic@gmail.com> Date: Sat, 18 Apr 2026 16:31:17 -0400 Subject: [PATCH 29/70] refactor(operator): resolve container via annotation and tidy deployment template - findOpenFGAContainer now reads the openfga.dev/container-name annotation emitted by the chart, and returns an error when the target container is missing instead of silently falling back to the first container in the pod spec. - migration_controller surfaces that error to the reconciler instead of logging and skipping, so misconfigured Deployments are visible. - deployment.yaml emits the new container-name annotation, collapses the replica-preservation logic to a single branch (both previous branches already preserved existing replicas), and uses selectorLabels on the pod template to avoid chart-version churn in pod labels across upgrades. - values.yaml documents the openfga-operator subchart values passthrough and clarifies migration service account behavior. --- charts/openfga/templates/deployment.yaml | 22 ++++----------- charts/openfga/values.yaml | 20 +++++++++++++ operator/internal/controller/helpers.go | 28 ++++++++++++------- .../controller/migration_controller.go | 10 +++---- 4 files changed, 48 insertions(+), 32 deletions(-) diff --git a/charts/openfga/templates/deployment.yaml b/charts/openfga/templates/deployment.yaml index a0da306f..e3d8601c 100644 --- a/charts/openfga/templates/deployment.yaml +++ b/charts/openfga/templates/deployment.yaml @@ -9,6 +9,7 @@ metadata: annotations: {{- if $hasOperatorAnnotations }} openfga.dev/migration-enabled: "true" + openfga.dev/container-name: "{{ .Chart.Name }}" openfga.dev/desired-replicas: '{{ ternary 1 .Values.replicaCount (eq .Values.datastore.engine "memory") }}' {{- if or .Values.migration.serviceAccount.create .Values.migration.serviceAccount.name }} openfga.dev/migration-service-account: '{{ include "openfga.migrationServiceAccountName" . }}' @@ -23,23 +24,10 @@ spec: {{- if .Values.autoscaling.enabled }} {{- fail "operator.enabled and autoscaling.enabled cannot both be true" }} {{- end }} - {{- /* On upgrade, preserve live replica count so existing pods keep serving (zero-downtime). - On fresh install, lookup returns empty — fall back to 0 so no pods start before migration. - See ADR-002 for rationale. */ -}} + {{- /* On upgrade: preserve live replicas (zero-downtime). On fresh install: lookup returns empty, fall back to 0. + OpenFGA gates readiness on MinimumSupportedDatastoreSchemaRevision — see ADR-002. */ -}} {{- $existing := (lookup "apps/v1" "Deployment" (include "openfga.namespace" .) (include "openfga.fullname" .)) }} - {{- $desiredImage := printf "%s:%s" .Values.image.repository (.Values.image.tag | default .Chart.AppVersion) }} - {{- $currentImage := "" }} - {{- range (($existing).spec).template.spec.containers }} - {{- if eq .name "openfga" }} - {{- $currentImage = .image }} - {{- end }} - {{- end }} - {{- if and $existing (eq $currentImage $desiredImage) }} - replicas: {{ $existing.spec.replicas }} - {{- else if $existing }} - {{- /* Image changed — preserve replicas; OpenFGA's built-in schema version check - (MinimumSupportedDatastoreSchemaRevision) causes readiness to fail until - migration completes, so old pods keep serving while new pods wait. */ -}} + {{- if and $existing (hasKey ($existing) "spec") }} replicas: {{ $existing.spec.replicas }} {{- else }} replicas: {{ ternary 1 0 (eq .Values.datastore.engine "memory") }} @@ -60,7 +48,7 @@ spec: prometheus.io/path: /metrics prometheus.io/port: "{{ (split ":" .Values.telemetry.metrics.addr)._1 }}" labels: - {{- include "openfga.labels" . | nindent 8 }} + {{- include "openfga.selectorLabels" . | nindent 8 }} {{- with .Values.podExtraLabels }} {{- toYaml . | nindent 8 }} {{- end }} diff --git a/charts/openfga/values.yaml b/charts/openfga/values.yaml index 24723c26..fcca5b09 100644 --- a/charts/openfga/values.yaml +++ b/charts/openfga/values.yaml @@ -391,6 +391,23 @@ extraObjects: [] operator: enabled: false +# -- Values passed to the openfga-operator subchart (when operator.enabled is true). +# See charts/openfga-operator/values.yaml for all available options. +openfga-operator: {} + # migrationJob: + # backoffLimit: 3 + # activeDeadlineSeconds: 300 + # ttlSecondsAfterFinished: 300 + # leaderElection: + # enabled: true + # watchNamespace: "" + # resources: + # requests: + # cpu: 10m + # memory: 64Mi + # limits: + # memory: 128Mi + # -- migration controls operator-driven migration behavior. # Only used when operator.enabled is true. migration: @@ -398,8 +415,11 @@ migration: enabled: true serviceAccount: # -- Create a dedicated service account for migration Jobs. + # The migration Job inherits env vars (including secretKeyRef) from the OpenFGA container. + # If your datastore secret has RBAC restrictions, ensure this service account can read it. create: true # -- Annotations to add to the migration service account. + # Use this to attach cloud IAM roles (e.g., eks.amazonaws.com/role-arn) for DDL permissions. annotations: {} # -- The name of the migration service account. # If not set and create is true, defaults to {fullname}-migration. diff --git a/operator/internal/controller/helpers.go b/operator/internal/controller/helpers.go index 7eced335..da1c7179 100644 --- a/operator/internal/controller/helpers.go +++ b/operator/internal/controller/helpers.go @@ -25,6 +25,7 @@ const ( // Annotations set on the Deployment by the Helm chart / operator. AnnotationMigrationEnabled = "openfga.dev/migration-enabled" + AnnotationContainerName = "openfga.dev/container-name" AnnotationDesiredReplicas = "openfga.dev/desired-replicas" AnnotationMigrationServiceAccount = "openfga.dev/migration-service-account" AnnotationRetryAfter = "openfga.dev/migration-retry-after" @@ -71,18 +72,25 @@ func migrationJobName(deploymentName string) string { } // findOpenFGAContainer finds the OpenFGA container in the Deployment's pod spec. -// It looks for a container named "openfga" first, then falls back to the first container. -func findOpenFGAContainer(deployment *appsv1.Deployment) *corev1.Container { - for i := range deployment.Spec.Template.Spec.Containers { - if deployment.Spec.Template.Spec.Containers[i].Name == "openfga" { - return &deployment.Spec.Template.Spec.Containers[i] - } +// It checks the openfga.dev/container-name annotation first, then looks for a +// container named "openfga". Returns an error if no containers exist or the +// target container is not found. +func findOpenFGAContainer(deployment *appsv1.Deployment) (*corev1.Container, error) { + containers := deployment.Spec.Template.Spec.Containers + if len(containers) == 0 { + return nil, fmt.Errorf("deployment %s/%s has no containers", deployment.Namespace, deployment.Name) } - // Fallback: use the first container (for charts that don't name it "openfga"). - if len(deployment.Spec.Template.Spec.Containers) > 0 { - return &deployment.Spec.Template.Spec.Containers[0] + + targetName := deployment.Annotations[AnnotationContainerName] + if targetName == "" { + targetName = "openfga" } - return nil + for i := range containers { + if containers[i].Name == targetName { + return &containers[i], nil + } + } + return nil, fmt.Errorf("container %q not found in deployment %s/%s", targetName, deployment.Namespace, deployment.Name) } // buildMigrationJob constructs a migration Job for the given Deployment. diff --git a/operator/internal/controller/migration_controller.go b/operator/internal/controller/migration_controller.go index 6152df23..2812a6f0 100644 --- a/operator/internal/controller/migration_controller.go +++ b/operator/internal/controller/migration_controller.go @@ -54,10 +54,10 @@ func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) ( } // 3. Find the OpenFGA container and extract the desired version. - mainContainer := findOpenFGAContainer(deployment) - if mainContainer == nil { - logger.Info("deployment has no containers, skipping") - return ctrl.Result{}, nil + mainContainer, err := findOpenFGAContainer(deployment) + if err != nil { + logger.Error(err, "unable to find OpenFGA container") + return ctrl.Result{}, err } desiredVersion := extractImageTag(mainContainer.Image) @@ -73,7 +73,7 @@ func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) ( // 4. Check current migration status from ConfigMap. configMap := &corev1.ConfigMap{} cmName := migrationConfigMapName(req.Name) - err := r.Get(ctx, types.NamespacedName{Name: cmName, Namespace: req.Namespace}, configMap) + err = r.Get(ctx, types.NamespacedName{Name: cmName, Namespace: req.Namespace}, configMap) currentVersion := "" if err == nil { From 1e5ccdd525a33a5d0968237a1b5fb1e7963f43b8 Mon Sep 17 00:00:00 2001 From: Ed Milic <edmilic@gmail.com> Date: Sun, 19 Apr 2026 05:23:22 -0400 Subject: [PATCH 30/70] chore: remove ADRs not relevant to this PR --- .../003-declarative-store-lifecycle-crds.md | 199 ------------------ docs/adr/004-operator-deployment-model.md | 164 --------------- 2 files changed, 363 deletions(-) delete mode 100644 docs/adr/003-declarative-store-lifecycle-crds.md delete mode 100644 docs/adr/004-operator-deployment-model.md diff --git a/docs/adr/003-declarative-store-lifecycle-crds.md b/docs/adr/003-declarative-store-lifecycle-crds.md deleted file mode 100644 index a54ee44b..00000000 --- a/docs/adr/003-declarative-store-lifecycle-crds.md +++ /dev/null @@ -1,199 +0,0 @@ -# ADR-003: Declarative Store Lifecycle Management via CRDs - -- **Status:** Proposed -- **Date:** 2026-04-06 -- **Deciders:** OpenFGA Helm Charts maintainers -- **Related ADR:** [ADR-001](001-adopt-openfga-operator.md) - -## Context - -OpenFGA is an authorization service. After deploying the server, teams must perform several runtime operations to make it usable: - -1. **Create a store** — a logical container for authorization data -2. **Write an authorization model** — the DSL that defines types, relations, and permissions -3. **Write tuples** — the relationship data that the model operates on (e.g., "user:anne is owner of document:budget") - -Today, these operations happen outside Kubernetes — through the OpenFGA API, CLI (`fga`), or custom scripts in CI pipelines. There is no declarative, Kubernetes-native way to manage them. - -This creates several problems: - -- **No GitOps for authorization config** — authorization models live in scripts or API calls, not in version-controlled manifests that ArgoCD/FluxCD sync. -- **No drift detection** — if someone modifies a model or tuple via the API, there's no controller to detect and reconcile the change. -- **No cross-team ownership** — each team that uses OpenFGA must build their own tooling to manage stores and models. There's no standard pattern. -- **Manual coordination** — deploying a new version of an application that needs a model change requires coordinating the Helm upgrade with a separate model push. - -### Alternatives Considered - -**A. CLI wrapper in CI pipelines** - -Use the `fga` CLI in a CI/CD step after `helm upgrade` to create stores, push models, and write tuples. - -*Pros:* No new Kubernetes components. Works with any CI system. -*Cons:* Imperative, not declarative. No drift detection. Each team builds their own pipeline. Model changes are not atomic with deployments. No visibility in Kubernetes tooling. - -**B. Helm post-install hook Job** - -Add a Helm hook Job that runs `fga` CLI commands after installation. - -*Pros:* Stays within the Helm ecosystem. -*Cons:* Helm hooks are the exact problem we're solving in ADR-002. Same ArgoCD/FluxCD incompatibilities. Hook Jobs are fire-and-forget with no reconciliation. - -**C. CRDs managed by the operator (selected)** - -Expose `FGAStore`, `FGAModel`, and `FGATuples` as Custom Resource Definitions. The operator watches these resources and reconciles them against the OpenFGA API. - -*Pros:* Fully declarative. GitOps-native. Continuous reconciliation. Standard Kubernetes patterns. Teams own their auth config as manifests. -*Cons:* Requires the operator (ADR-001). CRD design and reconciliation logic add development scope. Tuple reconciliation is complex. - -## Decision - -Introduce three CRDs, built in stages after the migration handling (ADR-002) is complete: - -### Stage 2: FGAStore - -```yaml -apiVersion: openfga.dev/v1alpha1 -kind: FGAStore -metadata: - name: my-app - namespace: my-team -spec: - # Reference to the OpenFGA instance - openfgaRef: - url: openfga.openfga-system.svc:8081 - credentialsRef: - name: openfga-api-credentials # Secret with API key or client credentials - # Store display name - name: "my-app-store" -status: - storeId: "01HXYZ..." - ready: true - conditions: - - type: Ready - status: "True" - lastTransitionTime: "2026-04-06T12:00:00Z" -``` - -**Controller behavior:** -- On create: call `CreateStore` API, store the returned store ID in `.status.storeId` -- On delete: call `DeleteStore` API (with finalizer to ensure cleanup) -- Idempotent: if a store with the same name exists, adopt it rather than creating a duplicate -- Status: set `Ready` condition when store is confirmed to exist - -### Stage 3: FGAModel - -```yaml -apiVersion: openfga.dev/v1alpha1 -kind: FGAModel -metadata: - name: my-app-model - namespace: my-team -spec: - storeRef: - name: my-app # References an FGAStore in the same namespace - model: | - model - schema 1.1 - type user - type organization - relations - define member: [user] - define admin: [user] - type document - relations - define reader: [user, organization#member] - define writer: [user, organization#admin] - define owner: [user] -status: - modelId: "01HABC..." - ready: true - lastWrittenHash: "sha256:a1b2c3..." # Hash of the model DSL to detect changes - conditions: - - type: Ready - status: "True" - - type: InSync - status: "True" -``` - -**Controller behavior:** -- On create/update: hash the model DSL. If hash differs from `.status.lastWrittenHash`, call `WriteAuthorizationModel` API -- Store the returned model ID in `.status.modelId` -- Model writes are append-only in OpenFGA (each write creates a new version), so this is safe -- Validation: optionally validate DSL syntax before calling the API (fail-fast with a clear error condition) -- The controller does NOT delete old model versions — OpenFGA retains model history - -### Stage 4: FGATuples - -```yaml -apiVersion: openfga.dev/v1alpha1 -kind: FGATuples -metadata: - name: my-app-base-tuples - namespace: my-team -spec: - storeRef: - name: my-app - tuples: - - user: "user:anne" - relation: "owner" - object: "document:budget" - - user: "team:engineering#member" - relation: "reader" - object: "folder:engineering-docs" - - user: "organization:acme#admin" - relation: "writer" - object: "folder:engineering-docs" -status: - writtenCount: 3 - ready: true - lastReconciled: "2026-04-06T12:00:00Z" - conditions: - - type: Ready - status: "True" - - type: InSync - status: "True" -``` - -**Controller behavior:** -- Maintain an **ownership model** — the controller tracks which tuples it wrote (via annotations or a status field). It only manages tuples it owns, never deleting tuples written by the application at runtime. -- On reconciliation: diff the desired tuples (from spec) against owned tuples in the store - - Tuples in spec but not in store → write them - - Tuples in store (owned) but not in spec → delete them - - Tuples in store but not owned → leave them alone -- Pagination: handle large tuple sets that exceed API response limits -- Batching: use `Write` API with batch operations to minimize API calls - -**Scope limitation:** `FGATuples` is intended for **base/static tuples** — organizational structure, role assignments, resource hierarchies. It is NOT intended to replace application-level tuple writes for dynamic data (e.g., per-request access grants). The ownership model ensures these two concerns don't interfere. - -### CRD Design Principles - -1. **Namespace-scoped** — all CRDs are namespaced, allowing teams to manage their own stores/models/tuples in their namespace -2. **Reference-based** — `FGAModel` and `FGATuples` reference an `FGAStore` by name, not by store ID. The controller resolves the reference. -3. **Status-driven** — controllers report state via `.status.conditions` following Kubernetes conventions (`Ready`, `InSync`, error conditions) -4. **Finalizers for cleanup** — `FGAStore` uses a finalizer to ensure the store is deleted from OpenFGA when the CR is deleted -5. **Idempotent** — all operations are safe to retry. Re-running reconciliation produces the same result. -6. **`v1alpha1` API version** — signals that the CRD schema may change. We will promote to `v1beta1` and `v1` as the design stabilizes. - -## Consequences - -### Positive - -- **GitOps-native authorization management** — stores, models, and tuples are Kubernetes resources that ArgoCD/FluxCD sync from Git -- **Drift detection and reconciliation** — the operator continuously ensures the actual state matches the declared state -- **Cross-team standardization** — every team uses the same CRDs, eliminating custom scripts and CI hacks -- **Atomic deployments** — a team can include `FGAModel` in their application's Helm chart; model updates deploy alongside code changes -- **Visibility** — `kubectl get fgastores`, `kubectl get fgamodels`, `kubectl describe fgatuples` provide instant visibility into authorization configuration -- **RBAC integration** — Kubernetes RBAC controls who can create/modify stores, models, and tuples per namespace - -### Negative - -- **Significant development scope** — three controllers, each with its own reconciliation logic, error handling, and tests -- **Tuple reconciliation complexity** — diffing and ownership tracking for tuples is the most complex piece; edge cases around partial failures, pagination, and large tuple sets -- **CRD upgrade burden** — CRD schema changes require careful migration; Helm does not upgrade CRDs automatically -- **API dependency** — the operator must be able to reach the OpenFGA API; network issues or API downtime affect reconciliation -- **Not suitable for all tuple management** — dynamic, application-driven tuples should still be written via the API, not CRDs. Users must understand this boundary. - -### Risks - -- **FGATuples at scale** — for stores with millions of tuples, the reconciliation diff could be expensive. The ownership model mitigates this (only diff owned tuples), but documentation must clearly state that `FGATuples` is for base/static data, not high-volume dynamic writes. -- **Multi-cluster** — if OpenFGA serves multiple clusters, CRDs in one cluster may conflict with CRDs in another pointing at the same store. This is out of scope for `v1alpha1` but should be considered for future versions. diff --git a/docs/adr/004-operator-deployment-model.md b/docs/adr/004-operator-deployment-model.md deleted file mode 100644 index a5a693c6..00000000 --- a/docs/adr/004-operator-deployment-model.md +++ /dev/null @@ -1,164 +0,0 @@ -# ADR-004: Operator Deployment as Helm Subchart Dependency - -- **Status:** Proposed -- **Date:** 2026-04-06 -- **Deciders:** OpenFGA Helm Charts maintainers -- **Related ADR:** [ADR-001](001-adopt-openfga-operator.md) - -## Context - -The OpenFGA Operator (ADR-001) needs a deployment model — how do users install it alongside or independent of the OpenFGA server? - -There are several established patterns in the Kubernetes ecosystem: - -### Alternatives Considered - -**A. Standalone operator chart (install separately)** - -Users install the operator chart first, then install the OpenFGA chart. The operator watches for OpenFGA Deployments across namespaces. - -*Example:* -```bash -helm install openfga-operator openfga/openfga-operator -n openfga-system -helm install openfga openfga/openfga -n my-namespace -``` - -*Pros:* Clean separation of concerns. One operator instance serves multiple OpenFGA installations. Follows the OLM/OperatorHub pattern. -*Cons:* Two install steps. Ordering dependency — operator must exist before the chart is useful. Users must manage two releases. Harder to get started. - -**B. Operator bundled in the main chart (single chart, always installed)** - -The operator Deployment, RBAC, and CRDs are templates in the main OpenFGA chart. No subchart. - -*Pros:* Simplest for users — one chart, one install. No dependency management. -*Cons:* Chart becomes larger and harder to maintain. Users who manage the operator separately (e.g., cluster-wide) can't disable it. CRDs are tied to the application chart's release cycle. Multiple OpenFGA installations in the same cluster would deploy multiple operator instances. - -**C. Operator as a conditional subchart dependency (selected)** - -The operator is a separate Helm chart (`openfga-operator`) that the main chart declares as a conditional dependency. Disabled by default for backward compatibility; users opt in with `operator.enabled: true`. - -*Example:* -```bash -# Everything in one command -helm install openfga openfga/openfga \ - --set datastore.engine=postgres \ - --set operator.enabled=true - -# Or, operator managed separately -helm install openfga-operator openfga/openfga-operator -n openfga-system -helm install openfga openfga/openfga \ - --set operator.enabled=false -``` - -*Pros:* Single install for most users. Operator chart has its own versioning. Users can disable for standalone management. Clean separation in code. -*Cons:* Subchart dependency adds some Chart.yaml complexity. CRDs still need special handling (Helm's `crds/` directory or a pre-install hook). - -**D. OLM (Operator Lifecycle Manager) only** - -Publish the operator to OperatorHub. Users install via OLM. - -*Pros:* Standard pattern for OpenShift. Handles CRD upgrades, operator upgrades, and RBAC. -*Cons:* OLM is not available on all clusters (not standard on EKS, GKE, AKS). Adds a dependency on OLM itself. Doesn't help Helm-only users. - -## Decision - -The operator will be distributed as a **conditional Helm subchart dependency** of the main OpenFGA chart. - -### Chart Structure - -```text -helm-charts/ -├── charts/ -│ ├── openfga/ # Main chart (existing) -│ │ ├── Chart.yaml # Declares openfga-operator as dependency -│ │ ├── values.yaml # operator.enabled: false (opt-in) -│ │ ├── templates/ -│ │ └── crds/ # Empty in Stage 1 -│ │ -│ └── openfga-operator/ # Operator subchart (new) -│ ├── Chart.yaml -│ ├── values.yaml -│ ├── templates/ -│ │ ├── deployment.yaml -│ │ ├── serviceaccount.yaml -│ │ ├── role.yaml -│ │ └── rolebinding.yaml -│ └── crds/ # CRDs added in Stages 2-4 -│ ├── fgastore.yaml -│ ├── fgamodel.yaml -│ └── fgatuples.yaml -``` - -### Dependency Declaration - -```yaml -# charts/openfga/Chart.yaml -dependencies: - - name: openfga-operator - version: "0.1.x" - repository: "file://../openfga-operator" - condition: operator.enabled -``` - -> **Note:** The `file://` reference is used because the operator subchart lives in the same -> monorepo. When the charts are published, consumers pulling from a registry will resolve the -> dependency automatically via the chart's packaging. - -### CRD Handling - -Helm has specific behavior around CRDs: - -1. **`crds/` directory** — CRDs placed here are installed on `helm install` but are **never upgraded or deleted** by Helm. This is safe but requires manual CRD upgrades. - -2. **Pre-install/pre-upgrade hook Job** — a Job that runs `kubectl apply -f` on CRD manifests before the main install/upgrade. This handles upgrades but reintroduces Helm hooks (the problem ADR-002 solves). - -3. **Static manifests applied separately** — CRDs are published as a standalone YAML file. Users run `kubectl apply -f` before `helm install`. This is the pattern used by cert-manager, Istio, and Prometheus Operator. - -**Decision:** Use the `crds/` directory in the operator subchart for initial installation. Publish CRD manifests as a standalone artifact for upgrades. Document both paths clearly. - -```bash -# First install — Helm installs CRDs automatically -helm install openfga openfga/openfga - -# CRD upgrades — applied manually (Helm won't upgrade them) -kubectl apply -f https://github.com/openfga/helm-charts/releases/download/v0.2.0/crds.yaml -``` - -### Installation Modes - -| Mode | Command | Use case | -|------|---------|----------| -| **Default** (no operator) | `helm install openfga openfga/openfga` | Backward compatible. Uses Helm hooks for migration. | -| **All-in-one** | `helm install openfga openfga/openfga --set operator.enabled=true` | Single install with operator-managed migrations. | -| **Operator standalone** | `helm install op openfga/openfga-operator -n openfga-system` | Cluster-wide operator serving multiple OpenFGA instances. | - -### Multi-Instance Considerations - -When multiple OpenFGA installations exist in the same cluster, each installation gets its own operator instance. The operator is **namespace-scoped** — it only watches resources in its own namespace (or the namespace specified via `--watch-namespace`). This ensures independent OpenFGA installations never interfere with each other. - -```yaml -# Operator values -operator: - watchNamespace: "" # empty = watch own namespace only (default) -``` - -## Consequences - -### Positive - -- **Single `helm install` for most users** — no ordering dependencies, no manual operator setup -- **Opt-out available** — `operator.enabled: false` for users who manage it separately or don't need it -- **Independent versioning** — operator chart has its own version; can be released on a different cadence than the main chart -- **Clean code separation** — operator code and templates are in their own chart directory -- **Namespace isolation** — each operator instance is scoped to its own namespace, so multiple OpenFGA installations coexist safely -- **Consistent with ecosystem** — this is the same pattern used by charts that depend on Bitnami PostgreSQL, Redis, etc. - -### Negative - -- **CRD upgrade complexity** — Helm does not upgrade CRDs; users must apply CRD manifests separately on operator upgrades -- **Multiple operators in all-in-one mode** — if a user installs OpenFGA in three namespaces, they get three operator pods (wasteful). Documentation should recommend standalone mode for multi-instance clusters. -- **Subchart value passing** — configuring the operator requires prefixed values (e.g., `openfga-operator.image.tag`), which is slightly less ergonomic than top-level values - -### Neutral - -- **OLM support is not excluded** — the operator can be published to OperatorHub in the future alongside the Helm distribution. The two are not mutually exclusive. From 97830774ce7cf6932fc1bbd9a1879453a76e269c Mon Sep 17 00:00:00 2001 From: Ed Milic <edmilic@gmail.com> Date: Sun, 19 Apr 2026 05:28:09 -0400 Subject: [PATCH 31/70] fix: clear retry-after annotation after Job creation --- .../controller/migration_controller.go | 17 ++-- .../controller/migration_controller_test.go | 89 ++++++++++++++++++- 2 files changed, 97 insertions(+), 9 deletions(-) diff --git a/operator/internal/controller/migration_controller.go b/operator/internal/controller/migration_controller.go index 2812a6f0..32757038 100644 --- a/operator/internal/controller/migration_controller.go +++ b/operator/internal/controller/migration_controller.go @@ -123,22 +123,23 @@ func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) ( r.ActiveDeadlineSeconds, r.TTLSecondsAfterFinished, ) - // Clear the retry-after annotation now that we're creating a new Job. - if _, hasRetry := deployment.Annotations[AnnotationRetryAfter]; hasRetry { - patch := client.MergeFrom(deployment.DeepCopy()) - delete(deployment.Annotations, AnnotationRetryAfter) - if patchErr := r.Patch(ctx, deployment, patch); patchErr != nil { - logger.Error(patchErr, "failed to clear retry-after annotation") - } - } if createErr := r.Create(ctx, job); createErr != nil { if apierrors.IsAlreadyExists(createErr) { // A concurrent reconcile already created the Job; requeue to pick it up. logger.V(1).Info("migration job already exists, will recheck", "job", jobName) return ctrl.Result{RequeueAfter: 5 * time.Second}, nil } + // Leave the retry-after annotation intact so the cooldown survives this failure. return ctrl.Result{}, fmt.Errorf("creating migration job: %w", createErr) } + // Clear the retry-after annotation now that the Job is created. + if _, hasRetry := deployment.Annotations[AnnotationRetryAfter]; hasRetry { + patch := client.MergeFrom(deployment.DeepCopy()) + delete(deployment.Annotations, AnnotationRetryAfter) + if patchErr := r.Patch(ctx, deployment, patch); patchErr != nil { + logger.Error(patchErr, "failed to clear retry-after annotation") + } + } logger.Info("created migration job", "job", jobName, "version", desiredVersion) return ctrl.Result{RequeueAfter: 5 * time.Second}, nil } else if err != nil { diff --git a/operator/internal/controller/migration_controller_test.go b/operator/internal/controller/migration_controller_test.go index a4d62a22..36037885 100644 --- a/operator/internal/controller/migration_controller_test.go +++ b/operator/internal/controller/migration_controller_test.go @@ -2,6 +2,7 @@ package controller import ( "context" + "fmt" "testing" "time" @@ -11,10 +12,12 @@ import ( metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" "k8s.io/apimachinery/pkg/runtime" "k8s.io/apimachinery/pkg/types" - "k8s.io/utils/ptr" clientgoscheme "k8s.io/client-go/kubernetes/scheme" + "k8s.io/utils/ptr" ctrl "sigs.k8s.io/controller-runtime" + "sigs.k8s.io/controller-runtime/pkg/client" "sigs.k8s.io/controller-runtime/pkg/client/fake" + "sigs.k8s.io/controller-runtime/pkg/client/interceptor" ) func newScheme() *runtime.Scheme { @@ -396,6 +399,90 @@ func TestReconcile_RetryAfterCooldown_SkipsJobCreation(t *testing.T) { } } +func TestReconcile_RetryAfterPersistsOnJobCreateFailure(t *testing.T) { + // Given: a Deployment with an elapsed retry-after annotation, and a client + // that fails Job creation with a non-AlreadyExists error. + dep := newTestDeployment("openfga", "default", "openfga/openfga:v1.14.0", 0) + dep.Annotations[AnnotationDesiredReplicas] = "3" + dep.Annotations[AnnotationRetryAfter] = time.Now().Add(-1 * time.Second).UTC().Format(time.RFC3339) + + scheme := newScheme() + c := fake.NewClientBuilder(). + WithScheme(scheme). + WithStatusSubresource(&appsv1.Deployment{}). + WithRuntimeObjects(dep). + WithInterceptorFuncs(interceptor.Funcs{ + Create: func(ctx context.Context, c client.WithWatch, obj client.Object, opts ...client.CreateOption) error { + if _, ok := obj.(*batchv1.Job); ok { + return fmt.Errorf("simulated transient API error") + } + return c.Create(ctx, obj, opts...) + }, + }). + Build() + r := &MigrationReconciler{ + Client: c, + BackoffLimit: DefaultBackoffLimit, + ActiveDeadlineSeconds: DefaultActiveDeadlineSeconds, + TTLSecondsAfterFinished: DefaultTTLSecondsAfterFinished, + } + + // When: reconciling. + _, err := r.Reconcile(context.Background(), ctrl.Request{ + NamespacedName: types.NamespacedName{Name: "openfga", Namespace: "default"}, + }) + + // Then: an error is returned and the retry-after annotation is preserved + // so the next reconcile honors the cooldown. + if err == nil { + t.Fatal("expected error from failed job creation") + } + + updated := &appsv1.Deployment{} + if getErr := r.Get(context.Background(), types.NamespacedName{ + Name: "openfga", Namespace: "default", + }, updated); getErr != nil { + t.Fatalf("getting deployment: %v", getErr) + } + if _, ok := updated.Annotations[AnnotationRetryAfter]; !ok { + t.Error("expected retry-after annotation to persist after Job creation failure") + } +} + +func TestReconcile_RetryAfterClearedAfterJobCreated(t *testing.T) { + // Given: a Deployment with an elapsed retry-after annotation. + dep := newTestDeployment("openfga", "default", "openfga/openfga:v1.14.0", 0) + dep.Annotations[AnnotationDesiredReplicas] = "3" + dep.Annotations[AnnotationRetryAfter] = time.Now().Add(-1 * time.Second).UTC().Format(time.RFC3339) + + r := newReconciler(dep) + + // When: reconciling. + if _, err := r.Reconcile(context.Background(), ctrl.Request{ + NamespacedName: types.NamespacedName{Name: "openfga", Namespace: "default"}, + }); err != nil { + t.Fatalf("unexpected error: %v", err) + } + + // Then: the Job exists and the retry-after annotation has been cleared. + job := &batchv1.Job{} + if getErr := r.Get(context.Background(), types.NamespacedName{ + Name: "openfga-migrate", Namespace: "default", + }, job); getErr != nil { + t.Fatalf("expected migration job to be created: %v", getErr) + } + + updated := &appsv1.Deployment{} + if getErr := r.Get(context.Background(), types.NamespacedName{ + Name: "openfga", Namespace: "default", + }, updated); getErr != nil { + t.Fatalf("getting deployment: %v", getErr) + } + if _, ok := updated.Annotations[AnnotationRetryAfter]; ok { + t.Error("expected retry-after annotation to be cleared after Job created") + } +} + func TestReconcile_MemoryDatastore_SkipsMigration(t *testing.T) { // Given: a Deployment using the memory datastore. dep := newTestDeployment("openfga", "default", "openfga/openfga:v1.14.0", 0) From 08c201a62488940cdd9151709d574390089b5d1a Mon Sep 17 00:00:00 2001 From: Ed Milic <edmilic@gmail.com> Date: Sun, 19 Apr 2026 05:33:42 -0400 Subject: [PATCH 32/70] fix: a migration Job without a version annotation or matching label is stale, its JobComplete would write the wrong version into the status ConfigMap --- .../controller/migration_controller.go | 13 ++-- .../controller/migration_controller_test.go | 68 +++++++++++++++++++ 2 files changed, 76 insertions(+), 5 deletions(-) diff --git a/operator/internal/controller/migration_controller.go b/operator/internal/controller/migration_controller.go index 32757038..ec718fb2 100644 --- a/operator/internal/controller/migration_controller.go +++ b/operator/internal/controller/migration_controller.go @@ -146,8 +146,11 @@ func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) ( return ctrl.Result{}, fmt.Errorf("getting migration job: %w", err) } - // 8. If the existing Job is for a different version, delete it and recreate. - // Check annotation first (supports digests > 63 chars), fall back to label. + // 8. If the existing Job is for a different (or unknown) version, delete it + // and recreate. Check annotation first (supports digests > 63 chars), fall + // back to label. A Job with neither marker is treated as stale: we cannot + // trust its outcome to represent the current desired version, so trusting + // JobComplete in step 9 would write a wrong version into the status ConfigMap. jobVersion := job.Annotations["openfga.dev/desired-version"] versionMatch := jobVersion == desiredVersion if jobVersion == "" { @@ -157,10 +160,10 @@ func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) ( sanitized = sanitized[:63] } jobVersion = job.Labels["app.kubernetes.io/version"] - versionMatch = jobVersion == sanitized + versionMatch = jobVersion != "" && jobVersion == sanitized } - if jobVersion != "" && !versionMatch { - logger.Info("existing migration job is for a different version, deleting", "jobVersion", jobVersion, "desiredVersion", desiredVersion) + if !versionMatch { + logger.Info("existing migration job is for a different or unknown version, deleting", "jobVersion", jobVersion, "desiredVersion", desiredVersion) propagation := metav1.DeletePropagationBackground if delErr := r.Delete(ctx, job, &client.DeleteOptions{ PropagationPolicy: &propagation, diff --git a/operator/internal/controller/migration_controller_test.go b/operator/internal/controller/migration_controller_test.go index 36037885..82996385 100644 --- a/operator/internal/controller/migration_controller_test.go +++ b/operator/internal/controller/migration_controller_test.go @@ -200,6 +200,9 @@ func TestReconcile_JobSucceeded_UpdatesConfigMapAndScalesUp(t *testing.T) { ObjectMeta: metav1.ObjectMeta{ Name: "openfga-migrate", Namespace: "default", + Annotations: map[string]string{ + "openfga.dev/desired-version": "v1.14.0", + }, OwnerReferences: []metav1.OwnerReference{ { APIVersion: "apps/v1", @@ -285,6 +288,9 @@ func TestReconcile_JobFailed_SetsRetryAnnotationAndRequeues(t *testing.T) { ObjectMeta: metav1.ObjectMeta{ Name: "openfga-migrate", Namespace: "default", + Annotations: map[string]string{ + "openfga.dev/desired-version": "v1.14.0", + }, OwnerReferences: []metav1.OwnerReference{ { APIVersion: "apps/v1", @@ -399,6 +405,65 @@ func TestReconcile_RetryAfterCooldown_SkipsJobCreation(t *testing.T) { } } +func TestReconcile_UnknownVersionJob_DeletedNotTrusted(t *testing.T) { + // Given: a Deployment desiring v1.14.0 and a JobComplete migration Job that + // carries no version annotation or label (e.g. left over from an older + // operator or created by a third-party tool). Trusting its outcome would + // write the wrong version into the migration-status ConfigMap. + dep := newTestDeployment("openfga", "default", "openfga/openfga:v1.14.0", 0) + dep.Annotations[AnnotationDesiredReplicas] = "3" + + job := &batchv1.Job{ + ObjectMeta: metav1.ObjectMeta{ + Name: "openfga-migrate", + Namespace: "default", + OwnerReferences: []metav1.OwnerReference{ + { + APIVersion: "apps/v1", + Kind: "Deployment", + Name: "openfga", + UID: "test-uid-123", + }, + }, + }, + Status: batchv1.JobStatus{ + Conditions: []batchv1.JobCondition{ + {Type: batchv1.JobComplete, Status: corev1.ConditionTrue}, + }, + }, + } + + r := newReconciler(dep, job) + + // When: reconciling. + result, err := r.Reconcile(context.Background(), ctrl.Request{ + NamespacedName: types.NamespacedName{Name: "openfga", Namespace: "default"}, + }) + + // Then: the Job is deleted and a requeue is scheduled; the ConfigMap is + // NOT created from the unknown-version Job's outcome. + if err != nil { + t.Fatalf("unexpected error: %v", err) + } + if result.RequeueAfter == 0 { + t.Error("expected requeue after deleting unknown-version job") + } + + deletedJob := &batchv1.Job{} + if getErr := r.Get(context.Background(), types.NamespacedName{ + Name: "openfga-migrate", Namespace: "default", + }, deletedJob); getErr == nil { + t.Error("expected unknown-version job to be deleted") + } + + cm := &corev1.ConfigMap{} + if getErr := r.Get(context.Background(), types.NamespacedName{ + Name: "openfga-migration-status", Namespace: "default", + }, cm); getErr == nil { + t.Errorf("expected no migration-status ConfigMap; got version=%q", cm.Data["version"]) + } +} + func TestReconcile_RetryAfterPersistsOnJobCreateFailure(t *testing.T) { // Given: a Deployment with an elapsed retry-after annotation, and a client // that fails Job creation with a non-AlreadyExists error. @@ -794,6 +859,9 @@ func TestReconcile_JobSucceeded_UpdatesExistingConfigMap(t *testing.T) { ObjectMeta: metav1.ObjectMeta{ Name: "openfga-migrate", Namespace: "default", + Annotations: map[string]string{ + "openfga.dev/desired-version": "v1.14.0", + }, OwnerReferences: []metav1.OwnerReference{ { APIVersion: "apps/v1", From 8cf9f704bfc062a8ceaa60709c52c93037a382d9 Mon Sep 17 00:00:00 2001 From: Ed Milic <edmilic@gmail.com> Date: Sun, 19 Apr 2026 05:42:52 -0400 Subject: [PATCH 33/70] fix(chart): restore full label set on pod template metadata MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit I switched the pod template labels from openfga.labels to openfga.selectorLabels, which would have stripped helm.sh/chart, commonLabels, component, version, managed-by, and part-of from running pods on upgrade — a breaking change for any tooling filtering on those labels. Add a helm-unittest regression guard for both operator on/off modes. --- charts/openfga/templates/deployment.yaml | 2 +- charts/openfga/tests/operator_mode_test.yaml | 33 ++++++++++++++++++++ 2 files changed, 34 insertions(+), 1 deletion(-) diff --git a/charts/openfga/templates/deployment.yaml b/charts/openfga/templates/deployment.yaml index e3d8601c..7318ddea 100644 --- a/charts/openfga/templates/deployment.yaml +++ b/charts/openfga/templates/deployment.yaml @@ -48,7 +48,7 @@ spec: prometheus.io/path: /metrics prometheus.io/port: "{{ (split ":" .Values.telemetry.metrics.addr)._1 }}" labels: - {{- include "openfga.selectorLabels" . | nindent 8 }} + {{- include "openfga.labels" . | nindent 8 }} {{- with .Values.podExtraLabels }} {{- toYaml . | nindent 8 }} {{- end }} diff --git a/charts/openfga/tests/operator_mode_test.yaml b/charts/openfga/tests/operator_mode_test.yaml index 2e509151..164d8bd1 100644 --- a/charts/openfga/tests/operator_mode_test.yaml +++ b/charts/openfga/tests/operator_mode_test.yaml @@ -135,3 +135,36 @@ tests: asserts: - isNotNull: path: spec.template.spec.initContainers + + # --- Pod template labels --- + # The pod template must carry the full common label set (helm.sh/chart, + # component, version, managed-by, part-of) — not just selectorLabels — + # so logging/monitoring tooling that filters on these labels keeps working + # across upgrades. Regression guard for the operator-migration branch. + - it: should include common labels on pod template metadata when operator is disabled + set: + operator.enabled: false + asserts: + - isNotEmpty: + path: spec.template.metadata.labels["helm.sh/chart"] + - equal: + path: spec.template.metadata.labels["app.kubernetes.io/component"] + value: authorization-controller + - equal: + path: spec.template.metadata.labels["app.kubernetes.io/part-of"] + value: openfga + + - it: should include common labels on pod template metadata when operator is enabled + set: + operator.enabled: true + migration.enabled: true + datastore.engine: postgres + asserts: + - isNotEmpty: + path: spec.template.metadata.labels["helm.sh/chart"] + - equal: + path: spec.template.metadata.labels["app.kubernetes.io/component"] + value: authorization-controller + - equal: + path: spec.template.metadata.labels["app.kubernetes.io/part-of"] + value: openfga From 9e3e1b39351c1118f63b7b181385c9a740f1f965 Mon Sep 17 00:00:00 2001 From: Ed Milic <edmilic@gmail.com> Date: Mon, 20 Apr 2026 05:43:02 -0400 Subject: [PATCH 34/70] =?UTF-8?q?ci:=20add=20operator-mode=20coverage=20an?= =?UTF-8?q?d=20v1.9.5=20=E2=86=92=20v1.14.1=20upgrade=20E2E?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - charts/openfga-operator/ci and charts/openfga/ci values files so chart-testing exercises both the standalone operator subchart and the parent chart with operator.enabled=true. - .github/ci/operator-postgres-values.yaml plus a dedicated workflow step that installs at v1.9.5 and upgrades to v1.14.1 (crossing the v1.10.0 !!REQUIRES MIGRATION!! boundary), asserting the migration ConfigMap and ready rollout at each step. --- .github/ci/operator-postgres-values.yaml | 77 +++++++++++++++++ .github/workflows/test.yml | 84 +++++++++++++++++++ .../openfga-operator/ci/default-values.yaml | 4 + charts/openfga/ci/operator-mode-values.yaml | 25 ++++++ 4 files changed, 190 insertions(+) create mode 100644 .github/ci/operator-postgres-values.yaml create mode 100644 charts/openfga-operator/ci/default-values.yaml create mode 100644 charts/openfga/ci/operator-mode-values.yaml diff --git a/.github/ci/operator-postgres-values.yaml b/.github/ci/operator-postgres-values.yaml new file mode 100644 index 00000000..4330d9ff --- /dev/null +++ b/.github/ci/operator-postgres-values.yaml @@ -0,0 +1,77 @@ +# E2E values consumed by the "operator + postgres E2E" step in test.yml. +# Not under charts/openfga/ci/ on purpose — chart-testing's helm-test runs +# a gRPC probe immediately after install, which would race the operator's +# scale-up. The dedicated workflow step waits for the migration ConfigMap +# and the scale-up explicitly, then verifies readiness. +replicaCount: 1 + +operator: + enabled: true + +migration: + enabled: true + +datastore: + engine: postgres + uriSecret: openfga-e2e-postgres-credentials + +openfga-operator: + image: + pullPolicy: Never + +extraObjects: + - apiVersion: v1 + kind: Secret + metadata: + name: openfga-e2e-postgres-credentials + stringData: + uri: "postgres://openfga:changeme@openfga-e2e-postgres:5432/openfga?sslmode=disable" + - apiVersion: apps/v1 + kind: Deployment + metadata: + name: openfga-e2e-postgres + spec: + replicas: 1 + selector: + matchLabels: + app: openfga-e2e-postgres + template: + metadata: + labels: + app: openfga-e2e-postgres + spec: + containers: + - name: postgres + image: postgres:17 + ports: + - containerPort: 5432 + env: + - name: POSTGRES_USER + value: openfga + - name: POSTGRES_PASSWORD + value: changeme + - name: POSTGRES_DB + value: openfga + - name: PGDATA + value: /var/lib/postgresql/data/pgdata + volumeMounts: + - name: data + mountPath: /var/lib/postgresql/data + readinessProbe: + exec: + command: ["pg_isready", "-U", "openfga", "-d", "openfga"] + initialDelaySeconds: 5 + periodSeconds: 5 + volumes: + - name: data + emptyDir: {} + - apiVersion: v1 + kind: Service + metadata: + name: openfga-e2e-postgres + spec: + selector: + app: openfga-e2e-postgres + ports: + - port: 5432 + targetPort: 5432 diff --git a/.github/workflows/test.yml b/.github/workflows/test.yml index bbec31d6..bc035c32 100644 --- a/.github/workflows/test.yml +++ b/.github/workflows/test.yml @@ -69,3 +69,87 @@ jobs: - name: Run chart-testing (install) if: steps.list-changed.outputs.changed == 'true' run: ct install --target-branch ${{ github.event.repository.default_branch }} + + - name: E2E test — operator-managed migration across schema boundary + id: e2e-operator + if: steps.list-changed.outputs.changed == 'true' + env: + NS: openfga-e2e + REL: openfga + # v1.9.5 → v1.14.1 crosses the v1.10.0 "!!REQUIRES MIGRATION!!" + # boundary (collation spec change in openfga/openfga#2661). + OLD_VER: v1.9.5 + NEW_VER: v1.14.1 + run: | + set -euo pipefail + kubectl create namespace "$NS" + helm dependency build charts/openfga + + echo "=== Phase 1: fresh install at ${OLD_VER} ===" + helm install "$REL" charts/openfga \ + --namespace "$NS" \ + --values .github/ci/operator-postgres-values.yaml \ + --set image.tag="${OLD_VER}" \ + --wait --timeout=3m + + # Operator pod must reach Ready (validates /readyz, RBAC, env vars). + kubectl wait deployment -n "$NS" \ + -l app.kubernetes.io/name=openfga-operator \ + --for=condition=Available=True --timeout=2m + + # Operator must run the migration Job and write ConfigMap at OLD_VER. + # Poll because kubectl wait --for=create requires kubectl >=1.31. + for i in $(seq 1 60); do + ver=$(kubectl get configmap "${REL}-migration-status" -n "$NS" \ + -o jsonpath='{.data.version}' 2>/dev/null || true) + if [ "$ver" = "${OLD_VER}" ]; then + echo "Phase 1: migration ConfigMap version=${ver}" + break + fi + sleep 3 + done + test "$ver" = "${OLD_VER}" + + # Operator must scale the openfga Deployment from 0 to 1 ready replica. + # condition=Available alone returns true at 0/0 before scale-up; + # readyReplicas=1 is the load-bearing signal. + kubectl wait deployment/"$REL" -n "$NS" \ + --for=jsonpath='{.status.readyReplicas}'=1 --timeout=3m + + echo "=== Phase 2: helm upgrade ${OLD_VER} → ${NEW_VER} ===" + helm upgrade "$REL" charts/openfga \ + --namespace "$NS" \ + --values .github/ci/operator-postgres-values.yaml \ + --set image.tag="${NEW_VER}" \ + --wait --timeout=3m + + # Operator must detect the version change, delete the stale Job, + # run a new migration, and update the ConfigMap to NEW_VER. + for i in $(seq 1 60); do + ver=$(kubectl get configmap "${REL}-migration-status" -n "$NS" \ + -o jsonpath='{.data.version}' 2>/dev/null || true) + if [ "$ver" = "${NEW_VER}" ]; then + echo "Phase 2: migration ConfigMap version=${ver}" + break + fi + sleep 3 + done + test "$ver" = "${NEW_VER}" + + # New pods must roll out at NEW_VER and become Ready. + kubectl wait deployment/"$REL" -n "$NS" \ + --for=jsonpath='{.status.readyReplicas}'=1 --timeout=3m + image=$(kubectl get deployment/"$REL" -n "$NS" \ + -o jsonpath='{.spec.template.spec.containers[0].image}') + echo "Phase 2 running image: $image" + echo "$image" | grep -q ":${NEW_VER}" + + - name: Dump operator E2E diagnostics on failure + if: failure() && steps.e2e-operator.conclusion == 'failure' + env: + NS: openfga-e2e + run: | + kubectl get all,configmap,job -n "$NS" -o wide || true + kubectl describe deployment -n "$NS" || true + kubectl logs -n "$NS" -l app.kubernetes.io/name=openfga-operator --tail=200 || true + kubectl logs -n "$NS" -l job-name --tail=200 || true diff --git a/charts/openfga-operator/ci/default-values.yaml b/charts/openfga-operator/ci/default-values.yaml new file mode 100644 index 00000000..93797cd5 --- /dev/null +++ b/charts/openfga-operator/ci/default-values.yaml @@ -0,0 +1,4 @@ +# Standalone install exercise for chart-testing. +# kind has the operator image preloaded, so skip the registry pull. +image: + pullPolicy: Never diff --git a/charts/openfga/ci/operator-mode-values.yaml b/charts/openfga/ci/operator-mode-values.yaml new file mode 100644 index 00000000..b85a6af2 --- /dev/null +++ b/charts/openfga/ci/operator-mode-values.yaml @@ -0,0 +1,25 @@ +# Exercises operator-managed mode end-to-end via chart-testing. +# +# The openfga-operator subchart auto-installs (conditional dependency on +# operator.enabled). With the memory datastore, the chart starts the +# Deployment at replicas=1 immediately, so `helm test` runs without racing +# the operator's reconcile loop. Migration is skipped (memory engine), but +# the rest of the wiring is exercised: subchart resolution, operator RBAC, +# pod/SA/annotation rendering, and the operator pod actually running and +# reconciling against the openfga Deployment in its release namespace. +# +# Postgres + operator (which exercises the migration Job path) is left to +# a follow-up E2E test — it requires waiting for the operator to scale the +# Deployment up before `helm test` runs the gRPC probe. +operator: + enabled: true + +migration: + enabled: true + +datastore: + engine: memory + +openfga-operator: + image: + pullPolicy: Never From 5a40ae2e7f72d10d5d5121830fb6096633f6b6ec Mon Sep 17 00:00:00 2001 From: Ed Milic <edmilic@gmail.com> Date: Mon, 20 Apr 2026 05:47:35 -0400 Subject: [PATCH 35/70] docs: update docs to reflect operator deployment changes --- docs/adr/002-operator-managed-migrations.md | 35 ++++++++++----------- docs/adr/README.md | 6 ++-- 2 files changed, 19 insertions(+), 22 deletions(-) diff --git a/docs/adr/002-operator-managed-migrations.md b/docs/adr/002-operator-managed-migrations.md index ad537992..1f0dc741 100644 --- a/docs/adr/002-operator-managed-migrations.md +++ b/docs/adr/002-operator-managed-migrations.md @@ -193,31 +193,30 @@ No hooks. No init containers. No `k8s-wait-for`. No downtime on upgrade. All res ### What Changes in the Helm Chart -**Removed:** +Nothing is deleted outright — every change is gated on `operator.enabled` so the legacy flow remains the default for backward compatibility. -| File/Section | Reason | -|--------------|--------| -| `templates/job.yaml` | Operator creates migration Jobs | -| `templates/rbac.yaml` | No init container polling Job status | -| `values.yaml`: `initContainer.repository`, `initContainer.tag` | `k8s-wait-for` eliminated | -| `values.yaml`: `datastore.migrationType` | Operator always uses Job internally | -| `values.yaml`: `datastore.waitForMigrations` | Operator handles ordering | -| `values.yaml`: `migrate.annotations` (hook annotations) | No Helm hooks | -| Deployment init containers for migration | Operator manages readiness via replica scaling | +**Gated on `operator.enabled: false` (legacy Helm-hook flow, rendered when the operator is disabled):** -**Added:** +| File/Section | Behavior when operator is enabled | +|--------------|-----------------------------------| +| `templates/job.yaml` | Skipped — operator creates migration Jobs dynamically | +| `templates/rbac.yaml` | Skipped — no init container needs to poll Job status | +| `values.yaml`: `initContainer.*` | Unused — `k8s-wait-for` not deployed | +| `values.yaml`: `datastore.migrationType`, `datastore.waitForMigrations` | Unused — operator always uses a Job and handles ordering | +| `values.yaml`: `migrate.annotations` | Unused — no Helm hooks | +| Deployment migration init containers | Skipped — operator manages readiness via replica scaling | + +**Added (active only when `operator.enabled: true`):** | File/Section | Purpose | |--------------|---------| -| `values.yaml`: `operator.enabled` | Toggle operator subchart | +| `values.yaml`: `operator.enabled` | Toggle the operator subchart | | `values.yaml`: `migration.serviceAccount.*` | Separate ServiceAccount for migration Jobs | -| `values.yaml`: `migration.timeout`, `backoffLimit`, `ttlSecondsAfterFinished` | Migration Job configuration | +| `values.yaml`: `migration.backoffLimit`, `activeDeadlineSeconds`, `ttlSecondsAfterFinished` | Migration Job configuration | | `templates/serviceaccount.yaml`: second SA | Migration ServiceAccount | -| `charts/openfga-operator/` | Operator subchart | - -**Preserved (backward compatible):** +| `charts/openfga-operator/` | Operator subchart (conditional dependency) | -When `operator.enabled: false`, the chart falls back to the current behavior — Helm hooks, `k8s-wait-for` init container, shared ServiceAccount. This allows gradual adoption. +Users on `operator.enabled: false` (the default) see identical rendered output to the pre-operator chart, so gradual adoption is possible with no forced migration. ## Consequences @@ -226,7 +225,7 @@ When `operator.enabled: false`, the chart falls back to the current behavior — - **All 6 migration issues resolved** — no Helm hooks means no ArgoCD/FluxCD/`--wait` incompatibility - **`k8s-wait-for` eliminated** — removes an unmaintained image with CVEs from the supply chain (#132, #144) - **Least-privilege enforced** — separate ServiceAccounts for migration (DDL) and runtime (CRUD) (#95) -- **Helm chart simplified** — 2 templates removed, init container logic removed, RBAC for job-watching removed +- **Runtime surface area reduced** — when `operator.enabled: true`, the legacy migration Job, init-container `k8s-wait-for` logic, and job-watching RBAC are skipped from the rendered manifest - **Migration is observable** — Job is a regular resource visible in all tools; ConfigMap records migration history; operator conditions surface errors - **Idempotent and crash-safe** — operator can restart at any point and resume correctly diff --git a/docs/adr/README.md b/docs/adr/README.md index 5f805122..d6b3445e 100644 --- a/docs/adr/README.md +++ b/docs/adr/README.md @@ -12,8 +12,6 @@ We follow the format described by [Michael Nygard](https://cognitect.com/blog/20 |-----|-------|--------|------| | [ADR-001](001-adopt-openfga-operator.md) | Adopt a Kubernetes Operator for OpenFGA Lifecycle Management | Proposed | 2026-04-06 | | [ADR-002](002-operator-managed-migrations.md) | Replace Helm Hook Migrations with Operator-Managed Migrations | Proposed | 2026-04-06 | -| [ADR-003](003-declarative-store-lifecycle-crds.md) | Declarative Store Lifecycle Management via CRDs | Proposed | 2026-04-06 | -| [ADR-004](004-operator-deployment-model.md) | Operator Deployment as Helm Subchart Dependency | Proposed | 2026-04-06 | --- @@ -67,9 +65,9 @@ When multiple ADRs are part of a single cohesive proposal — e.g., a foundation When doing this: -- **Explain the relationship in the PR description** — identify which ADR is the foundational decision and which are downstream. For example: "ADR-001 is the core decision to build an operator. ADR-002, 003, and 004 are downstream decisions about how the operator handles migrations, CRDs, and deployment." +- **Explain the relationship in the PR description** — identify which ADR is the foundational decision and which are downstream. For example: "ADR-001 is the core decision to build an operator. ADR-002 is a downstream decision about how the operator handles migrations." - **Each ADR can be accepted or rejected independently** — a reviewer might approve the foundational decision but push back on a downstream one. If that happens, split the PR: merge the accepted ADRs and keep the contested ones open for further discussion. -- **Keep each ADR self-contained** — even though they're in the same PR, each ADR should stand on its own. A reader should be able to understand ADR-003 without reading ADR-002 first (though they may reference each other). +- **Keep each ADR self-contained** — even though they're in the same PR, each ADR should stand on its own. A reader should be able to understand a downstream ADR without reading the foundational one first (though they may reference each other). ## How to Give Feedback on an ADR From c357a36388ca04790ca6b22a935e475e187cb83b Mon Sep 17 00:00:00 2001 From: Ed Milic <edmilic@gmail.com> Date: Mon, 20 Apr 2026 06:04:46 -0400 Subject: [PATCH 36/70] fix(operator): react to JobFailureTarget for fast failure detection MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The Job controller sets JobFailureTarget as soon as it decides a Job will fail (backoff limit reached, active deadline exceeded, etc.) — JobFailed only flips after pods finish terminating, which can take up to BackoffLimit × ActiveDeadlineSeconds. Previously the operator only watched JobFailed, so a broken migration took ~15 minutes (with chart defaults) before MigrationFailed appeared on the Deployment. Treat either condition as "failed" and add a regression test. --- .../controller/migration_controller.go | 7 +- .../controller/migration_controller_test.go | 66 +++++++++++++++++++ 2 files changed, 72 insertions(+), 1 deletion(-) diff --git a/operator/internal/controller/migration_controller.go b/operator/internal/controller/migration_controller.go index ec718fb2..ae096b63 100644 --- a/operator/internal/controller/migration_controller.go +++ b/operator/internal/controller/migration_controller.go @@ -197,7 +197,12 @@ func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) ( return ctrl.Result{}, nil } - if isJobConditionTrue(job, batchv1.JobFailed) { + // JobFailureTarget is set as soon as the Job controller decides the Job + // will fail (backoff limit reached, deadline exceeded, etc.); JobFailed + // only flips after pods finish terminating, which can take BackoffLimit × + // ActiveDeadlineSeconds. Treating either as "failed" surfaces the failure + // to users within seconds instead of minutes. + if isJobConditionTrue(job, batchv1.JobFailed) || isJobConditionTrue(job, batchv1.JobFailureTarget) { logger.Info("migration job failed, will delete and retry", "job", jobName, "version", desiredVersion) // Set condition so kubectl describe shows the failure. diff --git a/operator/internal/controller/migration_controller_test.go b/operator/internal/controller/migration_controller_test.go index 82996385..1bd2eb0b 100644 --- a/operator/internal/controller/migration_controller_test.go +++ b/operator/internal/controller/migration_controller_test.go @@ -372,6 +372,72 @@ func TestReconcile_JobFailed_SetsRetryAnnotationAndRequeues(t *testing.T) { } } +func TestReconcile_JobFailureTarget_TreatedAsFailed(t *testing.T) { + // Given: a Job with only JobFailureTarget=True (no JobFailed yet). The + // Job controller sets this as soon as it decides the Job will fail, + // before pods finish terminating and JobFailed is recorded. The operator + // should treat this as a failure to surface the error in seconds rather + // than waiting the full BackoffLimit × ActiveDeadlineSeconds. + dep := newTestDeployment("openfga", "default", "openfga/openfga:v1.14.0", 0) + dep.Annotations[AnnotationDesiredReplicas] = "3" + + job := &batchv1.Job{ + ObjectMeta: metav1.ObjectMeta{ + Name: "openfga-migrate", + Namespace: "default", + Annotations: map[string]string{ + "openfga.dev/desired-version": "v1.14.0", + }, + OwnerReferences: []metav1.OwnerReference{ + { + APIVersion: "apps/v1", + Kind: "Deployment", + Name: "openfga", + UID: "test-uid-123", + }, + }, + }, + Status: batchv1.JobStatus{ + Conditions: []batchv1.JobCondition{ + {Type: batchv1.JobFailureTarget, Status: corev1.ConditionTrue, Reason: "BackoffLimitExceeded"}, + }, + }, + } + + r := newReconciler(dep, job) + + result, err := r.Reconcile(context.Background(), ctrl.Request{ + NamespacedName: types.NamespacedName{Name: "openfga", Namespace: "default"}, + }) + if err != nil { + t.Fatalf("unexpected error: %v", err) + } + if result.RequeueAfter != 60*time.Second { + t.Errorf("expected 60s requeue, got %v", result.RequeueAfter) + } + + updated := &appsv1.Deployment{} + if getErr := r.Get(context.Background(), types.NamespacedName{ + Name: "openfga", Namespace: "default", + }, updated); getErr != nil { + t.Fatalf("getting deployment: %v", getErr) + } + + deletedJob := &batchv1.Job{} + if getErr := r.Get(context.Background(), types.NamespacedName{ + Name: "openfga-migrate", Namespace: "default", + }, deletedJob); getErr == nil { + t.Error("expected migration job to be deleted on JobFailureTarget") + } + if _, ok := updated.Annotations[AnnotationRetryAfter]; !ok { + t.Error("expected retry-after annotation to be set") + } + cond := findCondition(updated.Status.Conditions, "MigrationFailed") + if cond == nil || cond.Status != corev1.ConditionTrue { + t.Fatal("expected MigrationFailed condition True") + } +} + func TestReconcile_RetryAfterCooldown_SkipsJobCreation(t *testing.T) { // Given: a Deployment with a retry-after annotation in the future. dep := newTestDeployment("openfga", "default", "openfga/openfga:v1.14.0", 0) From be0c24b83f9ac6813c6937dd8088faa6c0b8c252 Mon Sep 17 00:00:00 2001 From: Ed Milic <edmilic@gmail.com> Date: Mon, 20 Apr 2026 06:08:23 -0400 Subject: [PATCH 37/70] chore(schema): reject unknown keys in operator and migration values MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Both the operator subchart and the parent chart's operator/migration blocks were missing additionalProperties: false, so typos like `migrationjob:` (lowercase), `enbaled: true`, or misplaced fields were silently ignored at install time. Add the guard to all well-defined object blocks — free-form blocks (podAnnotations, resources, securityContext, etc.) stay permissive since they pass through to pod spec. --- charts/openfga-operator/values.schema.json | 21 ++++++++++++++------- charts/openfga/values.schema.json | 13 ++++++++----- 2 files changed, 22 insertions(+), 12 deletions(-) diff --git a/charts/openfga-operator/values.schema.json b/charts/openfga-operator/values.schema.json index 470b265a..e74147c5 100644 --- a/charts/openfga-operator/values.schema.json +++ b/charts/openfga-operator/values.schema.json @@ -21,7 +21,8 @@ "type": "string" } }, - "required": ["repository"] + "required": ["repository"], + "additionalProperties": false }, "imagePullSecrets": { "type": "array", @@ -30,7 +31,8 @@ "properties": { "name": { "type": "string" } }, - "required": ["name"] + "required": ["name"], + "additionalProperties": false } }, "nameOverride": { "type": "string" }, @@ -42,7 +44,8 @@ "create": { "type": "boolean" }, "annotations": { "type": "object" }, "name": { "type": "string" } - } + }, + "additionalProperties": false }, "podAnnotations": { "type": "object" }, "podSecurityContext": { "type": "object" }, @@ -52,7 +55,8 @@ "type": "object", "properties": { "enabled": { "type": "boolean" } - } + }, + "additionalProperties": false }, "migrationJob": { "type": "object", @@ -69,7 +73,8 @@ "type": "integer", "minimum": 0 } - } + }, + "additionalProperties": false }, "resources": { "type": "object" }, "podDisruptionBudget": { @@ -88,7 +93,8 @@ { "type": "integer", "minimum": 0 } ] } - } + }, + "additionalProperties": false }, "nodeSelector": { "type": "object" }, "tolerations": { @@ -96,5 +102,6 @@ "items": { "type": "object" } }, "affinity": { "type": "object" } - } + }, + "additionalProperties": false } diff --git a/charts/openfga/values.schema.json b/charts/openfga/values.schema.json index 54e737b7..4fb19a27 100644 --- a/charts/openfga/values.schema.json +++ b/charts/openfga/values.schema.json @@ -1299,11 +1299,12 @@ "description": "Enable the openfga-operator subchart for operator-managed migrations", "default": false } - } + }, + "additionalProperties": false }, "openfga-operator": { "type": "object", - "description": "Values passed through to the openfga-operator subchart" + "description": "Values passed through to the openfga-operator subchart (validated by that chart's own schema)" }, "migration": { "type": "object", @@ -1332,12 +1333,14 @@ }, "name": { "type": "string", - "description": "The name of the migration service account. Defaults to {fullname}-migration.", + "description": "The name of the migration service account. Defaults to {fullname}-migration. Must be set explicitly when create=false and a dedicated migration SA is desired; leave empty to skip the annotation entirely.", "default": "" } - } + }, + "additionalProperties": false } - } + }, + "additionalProperties": false } }, "additionalProperties": false From 9a3b93354643de7df2ee9da35e68873a25845acd Mon Sep 17 00:00:00 2001 From: Ed Milic <edmilic@gmail.com> Date: Mon, 20 Apr 2026 06:11:42 -0400 Subject: [PATCH 38/70] ci(operator): build multi-arch on PRs, add immutable tag on main - Removes the push-to-main/dispatch gate on build-and-push so PRs verify the linux/arm64 build before merge. - On main pushes, adds an immutable :<version>-<sha> tag alongside the existing floating :<version> and :latest so consumers can pin a specific commit. --- .github/workflows/operator.yml | 32 ++++++++++++++++++++++---------- 1 file changed, 22 insertions(+), 10 deletions(-) diff --git a/.github/workflows/operator.yml b/.github/workflows/operator.yml index dc88c559..001d22f5 100644 --- a/.github/workflows/operator.yml +++ b/.github/workflows/operator.yml @@ -48,9 +48,6 @@ jobs: build-and-push: needs: test - if: >- - (github.event_name == 'push' && github.ref == 'refs/heads/main') || - (github.event_name == 'workflow_dispatch' && inputs.push_image) runs-on: ubuntu-latest permissions: contents: read @@ -68,31 +65,45 @@ jobs: echo "short_sha=${short_sha}" >> "$GITHUB_OUTPUT" echo "Operator version: ${version} (sha: ${short_sha})" - - name: Determine image tags + - name: Determine image tags and push policy id: tags run: | - if [[ "${{ github.ref }}" == "refs/heads/main" ]]; then - echo "tags=${{ env.IMAGE_NAME }}:${{ steps.version.outputs.version }},${{ env.IMAGE_NAME }}:latest" >> "$GITHUB_OUTPUT" - else - # Dev build — tag with version-sha to avoid clobbering release tags + if [[ "${{ github.event_name }}" == "push" && "${{ github.ref }}" == "refs/heads/main" ]]; then + # Main push: publish floating :<version> and :latest plus an + # immutable :<version>-<sha> so consumers pinning a specific + # commit have a stable reference. + echo "tags=${{ env.IMAGE_NAME }}:${{ steps.version.outputs.version }},${{ env.IMAGE_NAME }}:latest,${{ env.IMAGE_NAME }}:${{ steps.version.outputs.version }}-${{ steps.version.outputs.short_sha }}" >> "$GITHUB_OUTPUT" + echo "push=true" >> "$GITHUB_OUTPUT" + elif [[ "${{ github.event_name }}" == "workflow_dispatch" && "${{ inputs.push_image }}" == "true" ]]; then echo "tags=${{ env.IMAGE_NAME }}:${{ steps.version.outputs.version }}-${{ steps.version.outputs.short_sha }}" >> "$GITHUB_OUTPUT" + echo "push=true" >> "$GITHUB_OUTPUT" + else + # Pull request (or workflow_dispatch with push_image=false): + # build both platforms but don't publish — catches arm64-incompatible + # changes (build tags, syscalls, CGO) before they merge. + echo "tags=${{ env.IMAGE_NAME }}:pr-${{ steps.version.outputs.short_sha }}" >> "$GITHUB_OUTPUT" + echo "push=false" >> "$GITHUB_OUTPUT" fi + - name: Set up QEMU + uses: docker/setup-qemu-action@v3 + - name: Set up Docker Buildx uses: docker/setup-buildx-action@v3 - name: Login to GHCR + if: steps.tags.outputs.push == 'true' uses: docker/login-action@v4.1.0 with: registry: ghcr.io username: ${{ github.actor }} password: ${{ secrets.GITHUB_TOKEN }} - - name: Build and push + - name: Build and (conditionally) push uses: docker/build-push-action@v6 with: context: operator - push: true + push: ${{ steps.tags.outputs.push }} platforms: linux/amd64,linux/arm64 tags: ${{ steps.tags.outputs.tags }} cache-from: type=gha @@ -100,5 +111,6 @@ jobs: labels: | org.opencontainers.image.source=https://github.com/${{ github.repository }} org.opencontainers.image.version=${{ steps.version.outputs.version }} + org.opencontainers.image.revision=${{ github.sha }} org.opencontainers.image.title=openfga-operator org.opencontainers.image.description=OpenFGA Kubernetes operator for migration orchestration From d53b8e31213e4199053be21a0303580c63179e8f Mon Sep 17 00:00:00 2001 From: Ed Milic <edmilic@gmail.com> Date: Mon, 20 Apr 2026 06:15:04 -0400 Subject: [PATCH 39/70] docs(chart): explain 0-replica install in NOTES when operator is enabled When operator.enabled=true the workload starts at 0 replicas and only scales up after migration. If the operator pod is unhealthy this looks like a stuck install with no signal. Add NOTES output pointing at the operator deployment, migration Job, and MigrationFailed condition. --- charts/openfga/templates/NOTES.txt | 17 +++++++++++++++++ 1 file changed, 17 insertions(+) diff --git a/charts/openfga/templates/NOTES.txt b/charts/openfga/templates/NOTES.txt index 0048291e..628c3558 100644 --- a/charts/openfga/templates/NOTES.txt +++ b/charts/openfga/templates/NOTES.txt @@ -1,3 +1,20 @@ +{{- if and .Values.operator.enabled .Values.migration.enabled }} +NOTE: operator-managed migration is enabled. The OpenFGA Deployment starts at +0 replicas and is scaled up by the openfga-operator only after the migration +Job completes successfully. + +If pods don't appear within ~2 minutes, check the operator and the migration +Job: + + kubectl get deployment -A -l app.kubernetes.io/name=openfga-operator + kubectl logs -n {{ .Release.Namespace }} -l app.kubernetes.io/name=openfga-operator --tail=100 + kubectl get job/{{ include "openfga.fullname" . }}-migrate -n {{ .Release.Namespace }} -o yaml + kubectl describe deployment/{{ include "openfga.fullname" . }} -n {{ .Release.Namespace }} + +A `MigrationFailed` condition on the Deployment indicates the migration Job +failed; the operator will retry every 60s once the underlying issue is fixed. + +{{ end -}} 1. Get the application URL by running these commands: {{- if .Values.ingress.enabled }} {{- range $host := .Values.ingress.hosts }} From b945de26b71672029a099a590bac2c4546390b1a Mon Sep 17 00:00:00 2001 From: Ed Milic <edmilic@gmail.com> Date: Mon, 20 Apr 2026 06:19:30 -0400 Subject: [PATCH 40/70] fix: add missing global block to fix helm unit tests --- charts/openfga-operator/values.schema.json | 3 +++ 1 file changed, 3 insertions(+) diff --git a/charts/openfga-operator/values.schema.json b/charts/openfga-operator/values.schema.json index e74147c5..324465ff 100644 --- a/charts/openfga-operator/values.schema.json +++ b/charts/openfga-operator/values.schema.json @@ -2,6 +2,9 @@ "$schema": "https://json-schema.org/draft/2020-12/schema", "type": "object", "properties": { + "global": { + "type": "object" + }, "replicaCount": { "type": "integer", "minimum": 1 From 5ad48773541ae17118a10fdf995c078668f9b8ad Mon Sep 17 00:00:00 2001 From: Ed Milic <edmilic@gmail.com> Date: Mon, 20 Apr 2026 06:58:44 -0400 Subject: [PATCH 41/70] fix(operator): use multi-arch base image digests + Go cross-compile --- operator/Dockerfile | 16 +++++++++++----- 1 file changed, 11 insertions(+), 5 deletions(-) diff --git a/operator/Dockerfile b/operator/Dockerfile index 846414a1..034e22d0 100644 --- a/operator/Dockerfile +++ b/operator/Dockerfile @@ -1,5 +1,10 @@ -# pinned golang:1.26.2 linux/amd64 -FROM golang:1.26.2@sha256:b53c282df83967299380adbd6a2dc67e750a58217f39285d6240f6f80b19eaad AS builder +# pinned multi-arch index for golang:1.26.2 (linux/amd64, linux/arm64, ...) +FROM --platform=$BUILDPLATFORM golang:1.26.2@sha256:5f3787b7f902c07c7ec4f3aa91a301a3eda8133aa32661a3b3a3a86ab3a68a36 AS builder + +# buildx provides these automatically; declare so Go cross-compiles to the +# requested target instead of the build host's arch. +ARG TARGETOS +ARG TARGETARCH WORKDIR /workspace COPY go.mod go.sum ./ @@ -8,10 +13,11 @@ RUN go mod download COPY cmd/ cmd/ COPY internal/ internal/ -RUN CGO_ENABLED=0 GOOS=linux go build -ldflags="-s -w" -o /operator ./cmd/ +RUN CGO_ENABLED=0 GOOS=${TARGETOS} GOARCH=${TARGETARCH} \ + go build -ldflags="-s -w" -o /operator ./cmd/ -# pinned gcr.io/distroless/static:nonroot linux/amd64 -FROM gcr.io/distroless/static:nonroot@sha256:64c43684e6d2b581d1eb362ea47b6a4defee6a9cac5f7ebbda3daa67e8c9b8e6 +# pinned multi-arch index for gcr.io/distroless/static:nonroot +FROM gcr.io/distroless/static:nonroot@sha256:e3f945647ffb95b5839c07038d64f9811adf17308b9121d8a2b87b6a22a80a39 WORKDIR / COPY --from=builder /operator . USER 65532:65532 From 4c1e8a9cf7997a020985439652ebb48ce582aece Mon Sep 17 00:00:00 2001 From: Ed Milic <edmilic@gmail.com> Date: Mon, 20 Apr 2026 07:00:13 -0400 Subject: [PATCH 42/70] docs(operator-chart): clarify watchNamespace default --- charts/openfga-operator/values.yaml | 7 ++++++- 1 file changed, 6 insertions(+), 1 deletion(-) diff --git a/charts/openfga-operator/values.yaml b/charts/openfga-operator/values.yaml index db2747b4..921dcef7 100644 --- a/charts/openfga-operator/values.yaml +++ b/charts/openfga-operator/values.yaml @@ -39,7 +39,12 @@ securityContext: runAsUser: 65532 # -- Namespace to watch for OpenFGA Deployments. -# Leave empty to default to the release namespace. +# Leave empty to default to the operator pod's own namespace (read from +# the POD_NAMESPACE env var, set via the downward API). This usually +# equals the release namespace, but when `namespaceOverride` puts the +# operator in a different namespace than the release, the watch follows +# the pod — not the release. Set this explicitly to watch a specific +# namespace independent of where the operator runs. watchNamespace: "" leaderElection: From 4fb75513632287278abc91df2b170ba4b84bdc94 Mon Sep 17 00:00:00 2001 From: Ed Milic <edmilic@gmail.com> Date: Thu, 23 Apr 2026 04:37:32 -0400 Subject: [PATCH 43/70] docs: update ADRs to clarify operator migration status --- docs/adr/001-adopt-openfga-operator.md | 38 ++++++++++++++++----- docs/adr/002-operator-managed-migrations.md | 14 ++++---- 2 files changed, 37 insertions(+), 15 deletions(-) diff --git a/docs/adr/001-adopt-openfga-operator.md b/docs/adr/001-adopt-openfga-operator.md index ca845537..cf8c44c5 100644 --- a/docs/adr/001-adopt-openfga-operator.md +++ b/docs/adr/001-adopt-openfga-operator.md @@ -1,6 +1,6 @@ # ADR-001: Adopt a Kubernetes Operator for OpenFGA Lifecycle Management -- **Status:** Proposed +- **Status:** Accepted — Stage 1 implemented - **Date:** 2026-04-06 - **Deciders:** OpenFGA Helm Charts maintainers - **Related Issues:** #211, #107, #120, #100, #95, #126, #132, #143, #144 @@ -71,15 +71,36 @@ Development will follow a staged approach to deliver value incrementally: | 3 | `FGAModel` CRD | Declarative authorization model management | | 4 | `FGATuples` CRD | Declarative tuple management | +## Implementation Status + +Stage 1 has shipped on the `feat/operator-migration` branch. Stages 2-4 are planned but not yet implemented. + +### Delivered in Stage 1 + +- Operator Go project under `/operator/`, built with `controller-runtime` and kubebuilder scaffolding +- Operator packaged as a Helm subchart (`charts/openfga-operator/`) and wired into the main chart via a `condition: operator.enabled` dependency +- `operator.enabled` values toggle (default `false`) that gates all operator-managed behavior +- Migration reconciler (`migration_controller.go`) that orchestrates migration Jobs and gates Deployment readiness when the operator is enabled +- Separate migration ServiceAccount with IAM-annotation support (`openfga.migrationServiceAccountName` helper), created when the operator is enabled + +### Deferred to later stages + +- `FGAStore`, `FGAModel`, and `FGATuples` CRDs and their controllers — `charts/openfga-operator/crds/` is reserved but intentionally empty in Stage 1 +- Declarative store/model/tuple lifecycle management + +### Backward-compatibility path (deprecated) + +When `operator.enabled: false`, the chart still renders the legacy migration path: the Helm-hook migration Job, the `groundnuty/k8s-wait-for` init container, and the job-status RBAC. **This path is deprecated and will be removed in a future release** once the operator is the default and users have had time to migrate. It remains only to preserve backward compatibility during the transition. + ## Consequences ### Positive -- **Resolves all 6 migration issues** (#211, #107, #120, #100, #95, #126) and related dependency issues (#132, #144) -- **Eliminates `k8s-wait-for` dependency** — removes an unmaintained, CVE-carrying image from the supply chain -- **Enables GitOps-native authorization management** — stores, models, and tuples become declarative Kubernetes resources that ArgoCD/FluxCD can sync -- **Enforces least-privilege** — separate ServiceAccounts for migration (DDL) and runtime (CRUD) -- **Simplifies the Helm chart** — removes migration Job template, init container logic, RBAC for job-status-reading, and hook annotations +- **Resolves all 6 migration issues** (#211, #107, #120, #100, #95, #126) and related dependency issues (#132, #144) on the operator-enabled path +- **Removes `k8s-wait-for` from the operator-enabled path** — the unmaintained, CVE-carrying image is no longer used when `operator.enabled: true`, and will be removed from the chart entirely once the legacy path is retired +- **Enables GitOps-native authorization management** (planned, Stages 2-4) — stores, models, and tuples will become declarative Kubernetes resources that ArgoCD/FluxCD can sync +- **Enforces least-privilege** — separate ServiceAccounts for migration (DDL) and runtime (CRUD) on the operator-enabled path +- **Path to simplifying the Helm chart** — the migration Job template, init container logic, job-status RBAC, and hook annotations are conditionalized behind `operator.enabled: false` and scheduled for removal when the legacy path is retired - **Follows Kubernetes ecosystem conventions** — operators are the standard pattern for managing stateful application lifecycle ### Negative @@ -87,9 +108,10 @@ Development will follow a staged approach to deliver value incrementally: - **New component to maintain** — the operator is a full Go project with its own release cycle, CI, testing, and CVE surface - **Increased deployment footprint** — an additional pod running in the cluster (though resource requirements are minimal: ~50m CPU, ~64Mi memory) - **Learning curve** — contributors need to understand controller-runtime patterns to modify the operator -- **CRD management complexity** — Helm does not upgrade or delete CRDs; users may need to apply CRD manifests separately on operator upgrades +- **CRD management complexity** (applies once Stages 2-4 land) — Helm does not upgrade or delete CRDs; users may need to apply CRD manifests separately on operator upgrades +- **Two code paths during the transition** — the chart must maintain both the operator-enabled path and the deprecated legacy path until the latter is removed ### Neutral -- **Backward compatibility preserved** — the `operator.enabled: false` fallback maintains the existing Helm hook behavior for users who haven't migrated +- **Backward compatibility preserved during the transition** — `operator.enabled: false` keeps the existing Helm-hook behavior working for users who have not yet migrated, but this path is deprecated and slated for removal - **No change for memory-datastore users** — users running with `datastore.engine: memory` are unaffected (no migrations, no operator needed) diff --git a/docs/adr/002-operator-managed-migrations.md b/docs/adr/002-operator-managed-migrations.md index 1f0dc741..92a9abab 100644 --- a/docs/adr/002-operator-managed-migrations.md +++ b/docs/adr/002-operator-managed-migrations.md @@ -75,22 +75,22 @@ The operator runs a **migration controller** that reconciles the OpenFGA Deploym ``` ┌──────────────────────────────────────────────────────────┐ -│ Operator Reconciliation │ +│ Operator Reconciliation │ │ │ │ 1. Read Deployment → extract image tag (e.g. v1.14.0) │ -│ 2. Read ConfigMap/openfga-migration-status │ +│ 2. Read ConfigMap/openfga-migration-status │ │ └── "Last migrated version: v1.13.0" │ -│ 3. Versions differ → migration needed │ -│ 4. Create Job/openfga-migrate │ +│ 3. Versions differ → migration needed │ +│ 4. Create Job/openfga-migrate │ │ ├── ServiceAccount: openfga-migrator (DDL perms) │ │ ├── Image: openfga/openfga:v1.14.0 │ │ ├── Args: ["migrate"] │ │ └── ttlSecondsAfterFinished: 300 │ -│ 5. Watch Job until succeeded │ +│ 5. Watch Job until succeeded │ │ 6. Update ConfigMap → "version: v1.14.0" │ -│ 7. Ensure Deployment at desired replicas │ +│ 7. Ensure Deployment at desired replicas │ │ (fresh install: 0 → N; upgrade: already running) │ -│ 8. New pods pass readiness, serve requests │ +│ 8. New pods pass readiness, serve requests │ └──────────────────────────────────────────────────────────┘ ``` From e72ad4273c04c9625210e55f1216c8cf4744bcbd Mon Sep 17 00:00:00 2001 From: Anurag Bandyopadhyay <angbpy@gmail.com> Date: Tue, 22 Sep 2026 21:51:03 +0530 Subject: [PATCH 44/70] fix(operator): resolve migration reconcile churn and review findings - Make the MigrationFailed condition helpers idempotent and patch status only when the condition actually changes. The version-match path previously rewrote LastTransitionTime on every reconcile, which re-enqueued the Deployment and self-sustained continuous status writes once a Deployment carried a MigrationFailed condition. - Warn when the OpenFGA image is not pinned to an immutable tag or digest, since migrations are keyed on the image reference. - Drop the unused events RBAC grant; no EventRecorder is wired up. - Centralize the desired-version annotation and the managed-by/version label keys into constants instead of inline string literals. - Document the GitOps spec.replicas ownership caveat (ArgoCD ignoreDifferences / FluxCD field ownership) and the envFrom memory-datastore detection gap. - Add tests for condition idempotency and mutable-image detection. --- charts/openfga-operator/templates/role.yaml | 3 - operator/README.md | 12 +++ operator/internal/controller/helpers.go | 34 ++++++- .../controller/migration_controller.go | 90 ++++++++++++------ .../controller/migration_controller_test.go | 92 +++++++++++++++++++ 5 files changed, 194 insertions(+), 37 deletions(-) diff --git a/charts/openfga-operator/templates/role.yaml b/charts/openfga-operator/templates/role.yaml index dd17870b..29628ba3 100644 --- a/charts/openfga-operator/templates/role.yaml +++ b/charts/openfga-operator/templates/role.yaml @@ -21,6 +21,3 @@ rules: - apiGroups: ["coordination.k8s.io"] resources: ["leases"] verbs: ["get", "list", "watch", "create", "update"] - - apiGroups: [""] - resources: ["events"] - verbs: ["create", "patch"] diff --git a/operator/README.md b/operator/README.md index c9efbe8e..ea03c587 100644 --- a/operator/README.md +++ b/operator/README.md @@ -132,3 +132,15 @@ The operator reads these annotations from the OpenFGA Deployment: - **Mutable image tags:** The operator detects version changes by comparing the container image tag (or digest). If you deploy with a mutable tag like `latest` or reuse the same tag for different builds, the operator will not detect changes and will skip the migration. Use immutable tags (e.g., `v1.14.0`) or pin images by digest for reliable migration triggering. - **Migration-specific volumes:** The legacy Helm chart values `migrate.extraVolumes` and `migrate.extraVolumeMounts` have no effect in operator mode. The operator inherits volumes and mounts from the main Deployment pod spec. If you need additional volumes for migrations (e.g., CA bundles or TLS certs), add them to the top-level `extraVolumes` and `extraVolumeMounts` values instead. +- **`envFrom` datastore detection:** The memory-datastore check inspects only the explicit `env` entries on the container. If `OPENFGA_DATASTORE_ENGINE` is supplied via `envFrom` (a ConfigMap or Secret), the operator cannot read the value and will attempt a migration Job that a memory datastore does not need. The Helm chart sets this variable inline, so chart-managed installs are unaffected. +- **GitOps and `spec.replicas`:** The operator owns the Deployment's replica count — it scales to `0` for the duration of a migration and restores `openfga.dev/desired-replicas` afterward. Under a GitOps controller the chart's `lookup` of the live replica count returns empty at render time, so the rendered manifest carries `replicas: 0` (see [ADR-002](../docs/adr/002-operator-managed-migrations.md)). If the controller keeps syncing that field it will fight the operator over it, so exclude `spec.replicas` from GitOps reconciliation — the same treatment an HPA-managed Deployment needs: + - **ArgoCD** — add to the `Application`: + ```yaml + spec: + ignoreDifferences: + - group: apps + kind: Deployment + jsonPointers: + - /spec/replicas + ``` + - **FluxCD** — Flux applies server-side and honors field ownership, so drop `spec.replicas` from the Flux-applied manifest (e.g. a Kustomize patch removing `/spec/replicas`) and let the operator own it. diff --git a/operator/internal/controller/helpers.go b/operator/internal/controller/helpers.go index da1c7179..d5852333 100644 --- a/operator/internal/controller/helpers.go +++ b/operator/internal/controller/helpers.go @@ -3,6 +3,7 @@ package controller import ( "context" "fmt" + "regexp" "strconv" "strings" "time" @@ -23,10 +24,16 @@ const ( LabelPartOfValue = "openfga" LabelComponentValue = "authorization-controller" + // Labels set on operator-managed resources (migration Jobs, status ConfigMaps). + LabelManagedBy = "app.kubernetes.io/managed-by" + LabelVersion = "app.kubernetes.io/version" + LabelManagedByValue = "openfga-operator" + // Annotations set on the Deployment by the Helm chart / operator. AnnotationMigrationEnabled = "openfga.dev/migration-enabled" AnnotationContainerName = "openfga.dev/container-name" AnnotationDesiredReplicas = "openfga.dev/desired-replicas" + AnnotationDesiredVersion = "openfga.dev/desired-version" AnnotationMigrationServiceAccount = "openfga.dev/migration-service-account" AnnotationRetryAfter = "openfga.dev/migration-retry-after" @@ -36,6 +43,24 @@ const ( DefaultTTLSecondsAfterFinished int32 = 300 ) +// immutableTag matches a fully-qualified semantic version tag (with an optional +// leading "v" and optional pre-release/build suffix), which is treated as +// immutable by convention. +var immutableTag = regexp.MustCompile(`^v?\d+\.\d+\.\d+([-+][0-9A-Za-z.-]+)?$`) + +// isMutableImageReference reports whether an image reference is not pinned to an +// immutable identifier. Digest references (@sha256:...) are immutable, and a +// full semantic version tag is treated as immutable by convention. Everything +// else — "latest", a floating "v1.14", a bare name — is mutable: the same +// reference can resolve to different images over time, so the operator cannot +// tell that a rebuilt image needs a migration. +func isMutableImageReference(image string) bool { + if strings.Contains(image, "@") { + return false + } + return !immutableTag.MatchString(extractImageTag(image)) +} + // extractImageTag returns the tag portion of a container image reference. // For "openfga/openfga:v1.14.0" it returns "v1.14.0". // For "openfga/openfga@sha256:abc..." it returns the digest. @@ -122,11 +147,11 @@ func buildMigrationJob( Labels: map[string]string{ LabelPartOf: LabelPartOfValue, LabelComponent: "migration", - "app.kubernetes.io/managed-by": "openfga-operator", - "app.kubernetes.io/version": labelVersion, + LabelManagedBy: LabelManagedByValue, + LabelVersion: labelVersion, }, Annotations: map[string]string{ - "openfga.dev/desired-version": desiredVersion, + AnnotationDesiredVersion: desiredVersion, }, OwnerReferences: []metav1.OwnerReference{ { @@ -194,7 +219,7 @@ func updateMigrationStatus( Labels: map[string]string{ LabelPartOf: LabelPartOfValue, LabelComponent: "migration", - "app.kubernetes.io/managed-by": "openfga-operator", + LabelManagedBy: LabelManagedByValue, }, OwnerReferences: []metav1.OwnerReference{ { @@ -271,4 +296,3 @@ func ensureDeploymentScaled(ctx context.Context, c client.Client, deployment *ap } return false, nil } - diff --git a/operator/internal/controller/migration_controller.go b/operator/internal/controller/migration_controller.go index ae096b63..177ce0e7 100644 --- a/operator/internal/controller/migration_controller.go +++ b/operator/internal/controller/migration_controller.go @@ -70,6 +70,14 @@ func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) ( return ctrl.Result{}, nil } + // Migrations are keyed on the image reference. A mutable tag (e.g. "latest" + // or a floating "v1.14") can point at different images over time without the + // reference changing, so a rebuild that requires a migration will be missed. + if isMutableImageReference(mainContainer.Image) { + logger.Info("openfga image is not pinned to an immutable tag or digest; a migration may be silently skipped if the image changes without the tag changing", + "image", mainContainer.Image, "version", desiredVersion) + } + // 4. Check current migration status from ConfigMap. configMap := &corev1.ConfigMap{} cmName := migrationConfigMapName(req.Name) @@ -86,9 +94,10 @@ func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) ( if currentVersion == desiredVersion { logger.V(1).Info("migration up to date", "version", desiredVersion) statusPatch := client.MergeFrom(deployment.DeepCopy()) - clearMigrationFailedCondition(deployment) - if patchErr := r.Status().Patch(ctx, deployment, statusPatch); patchErr != nil { - logger.Error(patchErr, "failed to clear MigrationFailed condition") + if clearMigrationFailedCondition(deployment) { + if patchErr := r.Status().Patch(ctx, deployment, statusPatch); patchErr != nil { + logger.Error(patchErr, "failed to clear MigrationFailed condition") + } } if _, scaleErr := ensureDeploymentScaled(ctx, r.Client, deployment); scaleErr != nil { return ctrl.Result{}, scaleErr @@ -151,7 +160,7 @@ func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) ( // back to label. A Job with neither marker is treated as stale: we cannot // trust its outcome to represent the current desired version, so trusting // JobComplete in step 9 would write a wrong version into the status ConfigMap. - jobVersion := job.Annotations["openfga.dev/desired-version"] + jobVersion := job.Annotations[AnnotationDesiredVersion] versionMatch := jobVersion == desiredVersion if jobVersion == "" { // Label values have ":" replaced with "_", so sanitize desiredVersion for comparison. @@ -159,7 +168,7 @@ func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) ( if len(sanitized) > 63 { sanitized = sanitized[:63] } - jobVersion = job.Labels["app.kubernetes.io/version"] + jobVersion = job.Labels[LabelVersion] versionMatch = jobVersion != "" && jobVersion == sanitized } if !versionMatch { @@ -179,9 +188,10 @@ func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) ( // Clear MigrationFailed condition. statusPatch := client.MergeFrom(deployment.DeepCopy()) - clearMigrationFailedCondition(deployment) - if patchErr := r.Status().Patch(ctx, deployment, statusPatch); patchErr != nil { - logger.Error(patchErr, "failed to clear MigrationFailed condition") + if clearMigrationFailedCondition(deployment) { + if patchErr := r.Status().Patch(ctx, deployment, statusPatch); patchErr != nil { + logger.Error(patchErr, "failed to clear MigrationFailed condition") + } } // Update migration status ConfigMap. @@ -207,9 +217,10 @@ func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) ( // Set condition so kubectl describe shows the failure. statusPatch := client.MergeFrom(deployment.DeepCopy()) - setMigrationFailedCondition(deployment, desiredVersion) - if patchErr := r.Status().Patch(ctx, deployment, statusPatch); patchErr != nil { - logger.Error(patchErr, "failed to set MigrationFailed condition") + if setMigrationFailedCondition(deployment, desiredVersion) { + if patchErr := r.Status().Patch(ctx, deployment, statusPatch); patchErr != nil { + logger.Error(patchErr, "failed to set MigrationFailed condition") + } } // Persist a retry-after annotation so the cooldown is honored even @@ -269,37 +280,58 @@ func isMemoryDatastore(container *corev1.Container) bool { return false } -// setMigrationFailedCondition sets a MigrationFailed condition on the Deployment. -func setMigrationFailedCondition(deployment *appsv1.Deployment, version string) { - condition := appsv1.DeploymentCondition{ - Type: "MigrationFailed", - Status: corev1.ConditionTrue, - LastTransitionTime: metav1.Now(), - Reason: "MigrationJobFailed", - Message: fmt.Sprintf("Database migration failed for version %s. Check migration job logs.", version), - } - - // Replace existing MigrationFailed condition if present. +// setMigrationFailedCondition sets a MigrationFailed condition on the Deployment +// and reports whether anything actually changed, so callers can skip a no-op +// status write. LastTransitionTime only advances on a real status transition. +// +// NOTE: this writes a custom condition type onto a built-in Deployment's status. +// meta.SetStatusCondition cannot be used here because Deployment.Status.Conditions +// is []appsv1.DeploymentCondition, not []metav1.Condition. +func setMigrationFailedCondition(deployment *appsv1.Deployment, version string) bool { + message := fmt.Sprintf("Database migration failed for version %s. Check migration job logs.", version) for i, c := range deployment.Status.Conditions { if c.Type == "MigrationFailed" { - deployment.Status.Conditions[i] = condition - return + if c.Status == corev1.ConditionTrue && c.Reason == "MigrationJobFailed" && c.Message == message { + return false + } + if c.Status != corev1.ConditionTrue { + deployment.Status.Conditions[i].LastTransitionTime = metav1.Now() + } + deployment.Status.Conditions[i].Status = corev1.ConditionTrue + deployment.Status.Conditions[i].Reason = "MigrationJobFailed" + deployment.Status.Conditions[i].Message = message + return true } } - deployment.Status.Conditions = append(deployment.Status.Conditions, condition) + deployment.Status.Conditions = append(deployment.Status.Conditions, appsv1.DeploymentCondition{ + Type: "MigrationFailed", + Status: corev1.ConditionTrue, + LastTransitionTime: metav1.Now(), + Reason: "MigrationJobFailed", + Message: message, + }) + return true } -// clearMigrationFailedCondition removes or sets the MigrationFailed condition to False. -func clearMigrationFailedCondition(deployment *appsv1.Deployment) { +// clearMigrationFailedCondition sets an existing MigrationFailed condition to +// False and reports whether anything changed. When the condition is absent or +// already False it is a no-op — this is what stops the reconciler from patching +// status (and re-enqueueing the Deployment) on every reconcile of a healthy, +// version-matched Deployment. +func clearMigrationFailedCondition(deployment *appsv1.Deployment) bool { for i, c := range deployment.Status.Conditions { if c.Type == "MigrationFailed" { + if c.Status == corev1.ConditionFalse { + return false + } deployment.Status.Conditions[i].Status = corev1.ConditionFalse deployment.Status.Conditions[i].LastTransitionTime = metav1.Now() deployment.Status.Conditions[i].Reason = "MigrationSucceeded" deployment.Status.Conditions[i].Message = "Migration completed successfully." - return + return true } } + return false } // SetupWithManager sets up the controller with the Manager. @@ -322,7 +354,7 @@ func (r *MigrationReconciler) SetupWithManager(mgr ctrl.Manager) error { func(ctx context.Context, obj client.Object) []reconcile.Request { // Only watch ConfigMaps that are migration status ConfigMaps. if obj.GetLabels()[LabelPartOf] != LabelPartOfValue || - obj.GetLabels()["app.kubernetes.io/managed-by"] != "openfga-operator" { + obj.GetLabels()[LabelManagedBy] != LabelManagedByValue { return nil } // Map back to the owning Deployment. diff --git a/operator/internal/controller/migration_controller_test.go b/operator/internal/controller/migration_controller_test.go index 1bd2eb0b..b7ab5317 100644 --- a/operator/internal/controller/migration_controller_test.go +++ b/operator/internal/controller/migration_controller_test.go @@ -1102,3 +1102,95 @@ func TestExtractImageTag(t *testing.T) { }) } } + +func TestClearMigrationFailedConditionIdempotent(t *testing.T) { + // Absent condition: nothing to clear. + dep := &appsv1.Deployment{} + if clearMigrationFailedCondition(dep) { + t.Error("expected no change when the MigrationFailed condition is absent") + } + if len(dep.Status.Conditions) != 0 { + t.Errorf("expected no conditions to be added, got %d", len(dep.Status.Conditions)) + } + + // Condition present and True: clearing flips it to False (a real change). + dep.Status.Conditions = []appsv1.DeploymentCondition{{ + Type: "MigrationFailed", + Status: corev1.ConditionTrue, + }} + if !clearMigrationFailedCondition(dep) { + t.Error("expected a change when clearing a True MigrationFailed condition") + } + cond := findCondition(dep.Status.Conditions, "MigrationFailed") + if cond == nil || cond.Status != corev1.ConditionFalse { + t.Fatalf("expected MigrationFailed=False after clear, got %+v", cond) + } + transition := cond.LastTransitionTime + + // Already False: clearing again must be a no-op and must not advance + // LastTransitionTime — this is what stops the reconcile status write-churn. + if clearMigrationFailedCondition(dep) { + t.Error("expected no change when the MigrationFailed condition is already False") + } + cond = findCondition(dep.Status.Conditions, "MigrationFailed") + if !cond.LastTransitionTime.Equal(&transition) { + t.Error("LastTransitionTime must not change when the condition is already False") + } +} + +func TestSetMigrationFailedConditionIdempotent(t *testing.T) { + dep := &appsv1.Deployment{} + + // First set appends the condition. + if !setMigrationFailedCondition(dep, "v1.14.0") { + t.Error("expected a change when setting MigrationFailed on a fresh deployment") + } + cond := findCondition(dep.Status.Conditions, "MigrationFailed") + if cond == nil || cond.Status != corev1.ConditionTrue { + t.Fatalf("expected MigrationFailed=True, got %+v", cond) + } + transition := cond.LastTransitionTime + + // Re-setting for the same version is a no-op: no LastTransitionTime churn. + if setMigrationFailedCondition(dep, "v1.14.0") { + t.Error("expected no change when re-setting the same MigrationFailed condition") + } + cond = findCondition(dep.Status.Conditions, "MigrationFailed") + if !cond.LastTransitionTime.Equal(&transition) { + t.Error("LastTransitionTime must not change when the condition is unchanged") + } + + // A different version updates the message but does not transition status, + // so LastTransitionTime stays put. + if !setMigrationFailedCondition(dep, "v1.15.0") { + t.Error("expected a change when the failure message changes") + } + cond = findCondition(dep.Status.Conditions, "MigrationFailed") + if !cond.LastTransitionTime.Equal(&transition) { + t.Error("LastTransitionTime must not change without a status transition") + } +} + +func TestIsMutableImageReference(t *testing.T) { + tests := []struct { + image string + mutable bool + }{ + {"openfga/openfga:v1.14.0", false}, + {"openfga/openfga:1.14.0", false}, + {"openfga/openfga:v1.14.0-rc1", false}, + {"openfga/openfga@sha256:abcdef1234567890", false}, + {"openfga/openfga:latest", true}, + {"openfga/openfga:v1.14", true}, + {"openfga/openfga", true}, + {"registry.example.com:5000/openfga/openfga:v1.14.0", false}, + {"registry.example.com:5000/openfga/openfga:latest", true}, + } + for _, tt := range tests { + t.Run(tt.image, func(t *testing.T) { + if got := isMutableImageReference(tt.image); got != tt.mutable { + t.Errorf("isMutableImageReference(%q) = %v, want %v", tt.image, got, tt.mutable) + } + }) + } +} From 26a04dac32e099f7a8fc939a3e29f4375094900d Mon Sep 17 00:00:00 2001 From: Anurag Bandyopadhyay <angbpy@gmail.com> Date: Tue, 22 Sep 2026 21:59:35 +0530 Subject: [PATCH 45/70] chore: bump openfga chart to 0.4.0 for operator integration --- charts/openfga/Chart.yaml | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/charts/openfga/Chart.yaml b/charts/openfga/Chart.yaml index d03b93e5..94189434 100644 --- a/charts/openfga/Chart.yaml +++ b/charts/openfga/Chart.yaml @@ -3,7 +3,7 @@ name: openfga description: A Kubernetes Helm chart for the OpenFGA project. type: application -version: 0.3.13 +version: 0.4.0 appVersion: "v1.19.0" home: "https://openfga.github.io/helm-charts" From 3676f42fe3dfebce1117e7005672ebc52e42e5cf Mon Sep 17 00:00:00 2001 From: Siddhant Khare <siddhant@usegitai.com> Date: Tue, 22 Sep 2026 23:31:20 +0530 Subject: [PATCH 46/70] fix(operator): harden migration lifecycle --- .github/workflows/operator.yml | 4 + .github/workflows/test.yml | 1 + charts/openfga-operator/templates/NOTES.txt | 4 +- .../openfga-operator/templates/_helpers.tpl | 7 ++ charts/openfga-operator/templates/pdb.yaml | 4 +- charts/openfga-operator/templates/role.yaml | 2 +- .../templates/rolebinding.yaml | 2 +- charts/openfga-operator/tests/pdb_test.yaml | 25 ++++ .../tests/watch_namespace_role_test.yaml | 13 ++ .../watch_namespace_rolebinding_test.yaml | 16 +++ charts/openfga-operator/values.yaml | 3 +- charts/openfga/templates/NOTES.txt | 9 +- docs/adr/002-operator-managed-migrations.md | 12 +- operator/README.md | 14 +-- operator/cmd/main.go | 25 ++-- .../controller/migration_controller.go | 6 +- .../controller/migration_controller_test.go | 114 +++++++++++++++++- operator/tests/README.md | 2 +- 18 files changed, 221 insertions(+), 42 deletions(-) create mode 100644 charts/openfga-operator/tests/pdb_test.yaml create mode 100644 charts/openfga-operator/tests/watch_namespace_role_test.yaml create mode 100644 charts/openfga-operator/tests/watch_namespace_rolebinding_test.yaml diff --git a/.github/workflows/operator.yml b/.github/workflows/operator.yml index 001d22f5..a09638c9 100644 --- a/.github/workflows/operator.yml +++ b/.github/workflows/operator.yml @@ -42,6 +42,10 @@ jobs: working-directory: operator run: go test ./... -v + - name: Check formatting + working-directory: operator + run: test -z "$(gofmt -l .)" + - name: Run vet working-directory: operator run: go vet ./... diff --git a/.github/workflows/test.yml b/.github/workflows/test.yml index c30b30e5..7434585e 100644 --- a/.github/workflows/test.yml +++ b/.github/workflows/test.yml @@ -24,6 +24,7 @@ jobs: helm repo add bitnami-legacy https://raw.githubusercontent.com/bitnami/charts/archive-full-index/bitnami helm dependency build charts/openfga helm unittest charts/openfga + helm unittest charts/openfga-operator - uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7.0.0 with: diff --git a/charts/openfga-operator/templates/NOTES.txt b/charts/openfga-operator/templates/NOTES.txt index bcfcf801..47dba3bc 100644 --- a/charts/openfga-operator/templates/NOTES.txt +++ b/charts/openfga-operator/templates/NOTES.txt @@ -10,7 +10,7 @@ To view operator logs: kubectl logs --namespace {{ include "openfga-operator.namespace" . }} -l "app.kubernetes.io/name={{ include "openfga-operator.name" . }}" To check migration status: - kubectl get configmap -n {{ include "openfga-operator.namespace" . }} -l app.kubernetes.io/managed-by=openfga-operator + kubectl get configmap -n {{ include "openfga-operator.watchNamespace" . }} -l app.kubernetes.io/managed-by=openfga-operator To inspect migration jobs: - kubectl get jobs -n {{ include "openfga-operator.namespace" . }} -l app.kubernetes.io/part-of=openfga,app.kubernetes.io/component=migration + kubectl get jobs -n {{ include "openfga-operator.watchNamespace" . }} -l app.kubernetes.io/part-of=openfga,app.kubernetes.io/component=migration diff --git a/charts/openfga-operator/templates/_helpers.tpl b/charts/openfga-operator/templates/_helpers.tpl index f63057d6..dc7f512c 100644 --- a/charts/openfga-operator/templates/_helpers.tpl +++ b/charts/openfga-operator/templates/_helpers.tpl @@ -31,6 +31,13 @@ Allows overriding it for multi-namespace deployments in combined charts. {{- default .Release.Namespace .Values.namespaceOverride | trunc 63 | trimSuffix "-" -}} {{- end -}} +{{/* +Expand the namespace watched and managed by the operator. +*/}} +{{- define "openfga-operator.watchNamespace" -}} +{{- default (include "openfga-operator.namespace" .) .Values.watchNamespace | trunc 63 | trimSuffix "-" -}} +{{- end -}} + {{/* Create chart name and version as used by the chart label. */}} diff --git a/charts/openfga-operator/templates/pdb.yaml b/charts/openfga-operator/templates/pdb.yaml index 6c3514eb..9a068a3c 100644 --- a/charts/openfga-operator/templates/pdb.yaml +++ b/charts/openfga-operator/templates/pdb.yaml @@ -7,10 +7,10 @@ metadata: labels: {{- include "openfga-operator.labels" . | nindent 4 }} spec: - {{- if .Values.podDisruptionBudget.minAvailable }} + {{- if ne (toString .Values.podDisruptionBudget.minAvailable) "" }} minAvailable: {{ .Values.podDisruptionBudget.minAvailable }} {{- else }} - maxUnavailable: {{ .Values.podDisruptionBudget.maxUnavailable | default 1 }} + maxUnavailable: {{ .Values.podDisruptionBudget.maxUnavailable }} {{- end }} selector: matchLabels: diff --git a/charts/openfga-operator/templates/role.yaml b/charts/openfga-operator/templates/role.yaml index 29628ba3..fa2863a7 100644 --- a/charts/openfga-operator/templates/role.yaml +++ b/charts/openfga-operator/templates/role.yaml @@ -2,7 +2,7 @@ apiVersion: rbac.authorization.k8s.io/v1 kind: Role metadata: name: {{ include "openfga-operator.fullname" . }} - namespace: {{ include "openfga-operator.namespace" . }} + namespace: {{ include "openfga-operator.watchNamespace" . }} labels: {{- include "openfga-operator.labels" . | nindent 4 }} rules: diff --git a/charts/openfga-operator/templates/rolebinding.yaml b/charts/openfga-operator/templates/rolebinding.yaml index afacb98a..269bc3b6 100644 --- a/charts/openfga-operator/templates/rolebinding.yaml +++ b/charts/openfga-operator/templates/rolebinding.yaml @@ -2,7 +2,7 @@ apiVersion: rbac.authorization.k8s.io/v1 kind: RoleBinding metadata: name: {{ include "openfga-operator.fullname" . }} - namespace: {{ include "openfga-operator.namespace" . }} + namespace: {{ include "openfga-operator.watchNamespace" . }} labels: {{- include "openfga-operator.labels" . | nindent 4 }} roleRef: diff --git a/charts/openfga-operator/tests/pdb_test.yaml b/charts/openfga-operator/tests/pdb_test.yaml new file mode 100644 index 00000000..465beb34 --- /dev/null +++ b/charts/openfga-operator/tests/pdb_test.yaml @@ -0,0 +1,25 @@ +suite: operator PodDisruptionBudget +templates: + - templates/pdb.yaml +tests: + - it: should preserve zero minAvailable + set: + podDisruptionBudget.enabled: true + podDisruptionBudget.minAvailable: 0 + asserts: + - equal: + path: spec.minAvailable + value: 0 + - isNull: + path: spec.maxUnavailable + + - it: should preserve zero maxUnavailable + set: + podDisruptionBudget.enabled: true + podDisruptionBudget.maxUnavailable: 0 + asserts: + - equal: + path: spec.maxUnavailable + value: 0 + - isNull: + path: spec.minAvailable diff --git a/charts/openfga-operator/tests/watch_namespace_role_test.yaml b/charts/openfga-operator/tests/watch_namespace_role_test.yaml new file mode 100644 index 00000000..687925d3 --- /dev/null +++ b/charts/openfga-operator/tests/watch_namespace_role_test.yaml @@ -0,0 +1,13 @@ +suite: operator watch namespace Role +templates: + - templates/role.yaml +tests: + - it: should create managed-resource RBAC in the watch namespace + release: + namespace: operator-system + set: + watchNamespace: openfga-app + asserts: + - equal: + path: metadata.namespace + value: openfga-app diff --git a/charts/openfga-operator/tests/watch_namespace_rolebinding_test.yaml b/charts/openfga-operator/tests/watch_namespace_rolebinding_test.yaml new file mode 100644 index 00000000..258830e4 --- /dev/null +++ b/charts/openfga-operator/tests/watch_namespace_rolebinding_test.yaml @@ -0,0 +1,16 @@ +suite: operator watch namespace RoleBinding +templates: + - templates/rolebinding.yaml +tests: + - it: should bind the operator service account in the watch namespace + release: + namespace: operator-system + set: + watchNamespace: openfga-app + asserts: + - equal: + path: metadata.namespace + value: openfga-app + - equal: + path: subjects[0].namespace + value: operator-system diff --git a/charts/openfga-operator/values.yaml b/charts/openfga-operator/values.yaml index 921dcef7..ac211d4a 100644 --- a/charts/openfga-operator/values.yaml +++ b/charts/openfga-operator/values.yaml @@ -44,7 +44,8 @@ securityContext: # equals the release namespace, but when `namespaceOverride` puts the # operator in a different namespace than the release, the watch follows # the pod — not the release. Set this explicitly to watch a specific -# namespace independent of where the operator runs. +# namespace independent of where the operator runs. The target namespace +# must exist before installation so Helm can create the Role and RoleBinding. watchNamespace: "" leaderElection: diff --git a/charts/openfga/templates/NOTES.txt b/charts/openfga/templates/NOTES.txt index 628c3558..c8b16520 100644 --- a/charts/openfga/templates/NOTES.txt +++ b/charts/openfga/templates/NOTES.txt @@ -1,7 +1,8 @@ -{{- if and .Values.operator.enabled .Values.migration.enabled }} -NOTE: operator-managed migration is enabled. The OpenFGA Deployment starts at -0 replicas and is scaled up by the openfga-operator only after the migration -Job completes successfully. +{{- if and .Values.operator.enabled .Values.migration.enabled (has .Values.datastore.engine (list "postgres" "mysql")) }} +NOTE: operator-managed migration is enabled. A fresh installation starts the +OpenFGA Deployment at 0 replicas, then the operator scales it up after the +migration Job succeeds. An upgrade preserves the live replica count while the +migration runs. If pods don't appear within ~2 minutes, check the operator and the migration Job: diff --git a/docs/adr/002-operator-managed-migrations.md b/docs/adr/002-operator-managed-migrations.md index 92a9abab..282971b2 100644 --- a/docs/adr/002-operator-managed-migrations.md +++ b/docs/adr/002-operator-managed-migrations.md @@ -73,7 +73,7 @@ Replace the Helm hook migration Job and `k8s-wait-for` init container with **ope The operator runs a **migration controller** that reconciles the OpenFGA Deployment: -``` +```text ┌──────────────────────────────────────────────────────────┐ │ Operator Reconciliation │ │ │ @@ -125,7 +125,7 @@ The Job created by the operator has no Helm hook annotations. It is a standard K | Failure | Behavior | |---------|----------| -| Job fails | Operator sets `MigrationFailed` condition on Deployment. Does NOT scale up. User inspects Job logs. | +| Job fails | Operator sets `MigrationFailed` on the Deployment. A fresh installation remains at 0 replicas; an upgrade keeps its existing replicas. | | Job hangs | `activeDeadlineSeconds` (default 300s) kills it. Operator sees failure. | | Operator crashes | On restart, re-reads ConfigMap and Job status. Resumes from where it left off. | | Database unreachable | Job fails to connect. After exhausting `backoffLimit`, operator deletes the failed Job, sets a `retry-after` annotation, and recreates a fresh Job after a fixed 60-second cooldown. Cycle repeats until the database becomes available. | @@ -134,7 +134,7 @@ The Job created by the operator has no Helm hook annotations. It is a standard K **Before (Helm hooks):** -``` +```text helm install ├── Create ServiceAccount, RBAC, Secret, Service ├── Create Deployment (with wait-for-migration init container) @@ -151,7 +151,7 @@ Problems: ArgoCD skips step 4. FluxCD deletes Job in step 4. `--wait` deadlocks **After (operator-managed, fresh install):** -``` +```text helm install ├── Create ServiceAccount (runtime), ServiceAccount (migrator) ├── Create Secret, Service @@ -171,7 +171,7 @@ Operator starts: **After (operator-managed, upgrade with new image):** -``` +```text helm upgrade ├── lookup finds existing Deployment at 3 replicas → preserves replicas: 3 ├── Patches Deployment with new image tag @@ -212,7 +212,7 @@ Nothing is deleted outright — every change is gated on `operator.enabled` so t |--------------|---------| | `values.yaml`: `operator.enabled` | Toggle the operator subchart | | `values.yaml`: `migration.serviceAccount.*` | Separate ServiceAccount for migration Jobs | -| `values.yaml`: `migration.backoffLimit`, `activeDeadlineSeconds`, `ttlSecondsAfterFinished` | Migration Job configuration | +| `values.yaml`: `openfga-operator.migrationJob.*` | Migration Job backoff, deadline, and TTL configuration | | `templates/serviceaccount.yaml`: second SA | Migration ServiceAccount | | `charts/openfga-operator/` | Operator subchart (conditional dependency) | diff --git a/operator/README.md b/operator/README.md index ea03c587..6e1ab6fe 100644 --- a/operator/README.md +++ b/operator/README.md @@ -6,14 +6,14 @@ This is **Stage 1** of the operator — focused solely on migration orchestratio ## How It Works -1. The operator watches Deployments **in its own namespace** labeled `app.kubernetes.io/part-of: openfga` and `app.kubernetes.io/component: authorization-controller` +1. The operator watches Deployments in its configured namespace, which defaults to the operator pod's namespace, labeled `app.kubernetes.io/part-of: openfga` and `app.kubernetes.io/component: authorization-controller` 2. When a version change is detected (comparing the container image tag to the `{name}-migration-status` ConfigMap), the operator: - - Keeps the Deployment at 0 replicas + - Leaves existing replicas running during upgrades - Creates a migration Job running `openfga migrate` - Waits for the Job to complete - Updates the ConfigMap with the new version - - Scales the Deployment up to the desired replica count -3. On failure, a `MigrationFailed` condition is set on the Deployment and replicas stay at 0 + - Scales a fresh installation from 0 to the desired replica count +3. On failure, a `MigrationFailed` condition is set on the Deployment. Fresh installations remain at 0 replicas, while upgrades keep their existing replicas. ## Prerequisites @@ -87,7 +87,7 @@ See [`tests/README.md`](tests/README.md) for detailed verification steps and all ## Project Structure -``` +```text operator/ ├── cmd/ │ └── main.go # Entry point, manager setup @@ -109,7 +109,7 @@ The operator accepts the following flags: | Flag | Default | Description | |------|---------|-------------| | `--leader-elect` | `false` | Enable leader election so only one replica actively reconciles at a time. Required when running multiple operator replicas for high availability; standby pods wait for the leader's Lease to expire before taking over. Not needed for single-replica deployments. | -| `--watch-namespace` | `""` | Namespace to watch for OpenFGA Deployments. Defaults to the operator pod's own namespace (via `POD_NAMESPACE` env var). Each operator instance manages only its own namespace, so multiple independent OpenFGA installations can coexist safely. | +| `--watch-namespace` | `""` | Namespace to watch for OpenFGA Deployments. Defaults to the operator pod's own namespace (via `POD_NAMESPACE` env var). The chart binds namespaced RBAC in the configured watch namespace, so the operator may run in a different namespace when needed. | | `--metrics-bind-address` | `:8080` | Address the Prometheus metrics endpoint binds to. Change only if the default port conflicts with other containers in the pod. | | `--health-probe-bind-address` | `:8081` | Address the Kubernetes liveness and readiness probe endpoints bind to. Change only if the default port conflicts. | | `--backoff-limit` | `3` | Number of times a migration Job's pod can fail before the Job is considered failed. After hitting this limit the operator deletes the Job, sets a `MigrationFailed` condition on the Deployment, and retries after a 60-second cooldown. | @@ -133,7 +133,7 @@ The operator reads these annotations from the OpenFGA Deployment: - **Mutable image tags:** The operator detects version changes by comparing the container image tag (or digest). If you deploy with a mutable tag like `latest` or reuse the same tag for different builds, the operator will not detect changes and will skip the migration. Use immutable tags (e.g., `v1.14.0`) or pin images by digest for reliable migration triggering. - **Migration-specific volumes:** The legacy Helm chart values `migrate.extraVolumes` and `migrate.extraVolumeMounts` have no effect in operator mode. The operator inherits volumes and mounts from the main Deployment pod spec. If you need additional volumes for migrations (e.g., CA bundles or TLS certs), add them to the top-level `extraVolumes` and `extraVolumeMounts` values instead. - **`envFrom` datastore detection:** The memory-datastore check inspects only the explicit `env` entries on the container. If `OPENFGA_DATASTORE_ENGINE` is supplied via `envFrom` (a ConfigMap or Secret), the operator cannot read the value and will attempt a migration Job that a memory datastore does not need. The Helm chart sets this variable inline, so chart-managed installs are unaffected. -- **GitOps and `spec.replicas`:** The operator owns the Deployment's replica count — it scales to `0` for the duration of a migration and restores `openfga.dev/desired-replicas` afterward. Under a GitOps controller the chart's `lookup` of the live replica count returns empty at render time, so the rendered manifest carries `replicas: 0` (see [ADR-002](../docs/adr/002-operator-managed-migrations.md)). If the controller keeps syncing that field it will fight the operator over it, so exclude `spec.replicas` from GitOps reconciliation — the same treatment an HPA-managed Deployment needs: +- **GitOps and `spec.replicas`:** On a fresh installation, the chart renders `replicas: 0` and the operator sets `openfga.dev/desired-replicas` after migration. During an in-cluster Helm upgrade, `lookup` preserves the live replica count. GitOps renderers cannot perform that lookup, so their manifest still carries `replicas: 0` (see [ADR-002](../docs/adr/002-operator-managed-migrations.md)). Exclude `spec.replicas` from GitOps reconciliation so a later sync does not scale a healthy Deployment back to 0: - **ArgoCD** — add to the `Application`: ```yaml spec: diff --git a/operator/cmd/main.go b/operator/cmd/main.go index ac9bac64..daf01e8e 100644 --- a/operator/cmd/main.go +++ b/operator/cmd/main.go @@ -25,12 +25,12 @@ func init() { func main() { var ( - leaderElect bool - watchNamespace string - metricsAddr string - healthProbeAddr string - backoffLimit int - activeDeadline int + leaderElect bool + watchNamespace string + metricsAddr string + healthProbeAddr string + backoffLimit int + activeDeadline int ttlAfterFinished int ) @@ -82,12 +82,13 @@ func main() { } mgr, err := ctrl.NewManager(ctrl.GetConfigOrDie(), ctrl.Options{ - Scheme: scheme, - Metrics: metricsserver.Options{BindAddress: metricsAddr}, - HealthProbeBindAddress: healthProbeAddr, - LeaderElection: leaderElect, - LeaderElectionID: "openfga-operator-leader", - Cache: cacheOpts, + Scheme: scheme, + Metrics: metricsserver.Options{BindAddress: metricsAddr}, + HealthProbeBindAddress: healthProbeAddr, + LeaderElection: leaderElect, + LeaderElectionID: "openfga-operator-leader", + LeaderElectionNamespace: watchNamespace, + Cache: cacheOpts, }) if err != nil { logger.Error(err, "unable to create manager") diff --git a/operator/internal/controller/migration_controller.go b/operator/internal/controller/migration_controller.go index 177ce0e7..9131fa38 100644 --- a/operator/internal/controller/migration_controller.go +++ b/operator/internal/controller/migration_controller.go @@ -96,7 +96,7 @@ func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) ( statusPatch := client.MergeFrom(deployment.DeepCopy()) if clearMigrationFailedCondition(deployment) { if patchErr := r.Status().Patch(ctx, deployment, statusPatch); patchErr != nil { - logger.Error(patchErr, "failed to clear MigrationFailed condition") + return ctrl.Result{}, fmt.Errorf("clearing MigrationFailed condition: %w", patchErr) } } if _, scaleErr := ensureDeploymentScaled(ctx, r.Client, deployment); scaleErr != nil { @@ -190,7 +190,7 @@ func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) ( statusPatch := client.MergeFrom(deployment.DeepCopy()) if clearMigrationFailedCondition(deployment) { if patchErr := r.Status().Patch(ctx, deployment, statusPatch); patchErr != nil { - logger.Error(patchErr, "failed to clear MigrationFailed condition") + return ctrl.Result{}, fmt.Errorf("clearing MigrationFailed condition: %w", patchErr) } } @@ -219,7 +219,7 @@ func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) ( statusPatch := client.MergeFrom(deployment.DeepCopy()) if setMigrationFailedCondition(deployment, desiredVersion) { if patchErr := r.Status().Patch(ctx, deployment, statusPatch); patchErr != nil { - logger.Error(patchErr, "failed to set MigrationFailed condition") + return ctrl.Result{}, fmt.Errorf("setting MigrationFailed condition: %w", patchErr) } } diff --git a/operator/internal/controller/migration_controller_test.go b/operator/internal/controller/migration_controller_test.go index b7ab5317..08513cc0 100644 --- a/operator/internal/controller/migration_controller_test.go +++ b/operator/internal/controller/migration_controller_test.go @@ -83,6 +83,36 @@ func newReconciler(objects ...runtime.Object) *MigrationReconciler { } } +func newReconcilerWithStatusPatchError(objects ...runtime.Object) *MigrationReconciler { + scheme := newScheme() + clientBuilder := fake.NewClientBuilder().WithScheme(scheme). + WithStatusSubresource(&appsv1.Deployment{}) + for _, obj := range objects { + clientBuilder = clientBuilder.WithRuntimeObjects(obj) + } + c := clientBuilder.WithInterceptorFuncs(interceptor.Funcs{ + SubResourcePatch: func( + ctx context.Context, + c client.Client, + subResourceName string, + obj client.Object, + patch client.Patch, + opts ...client.SubResourcePatchOption, + ) error { + if subResourceName == "status" { + return fmt.Errorf("simulated status patch error") + } + return c.SubResource(subResourceName).Patch(ctx, obj, patch, opts...) + }, + }).Build() + return &MigrationReconciler{ + Client: c, + BackoffLimit: DefaultBackoffLimit, + ActiveDeadlineSeconds: DefaultActiveDeadlineSeconds, + TTLSecondsAfterFinished: DefaultTTLSecondsAfterFinished, + } +} + func findCondition(conditions []appsv1.DeploymentCondition, condType string) *appsv1.DeploymentCondition { for i := range conditions { if string(conditions[i].Type) == condType { @@ -182,6 +212,43 @@ func TestReconcile_VersionMatch_ScalesUp(t *testing.T) { } } +func TestReconcile_VersionMatch_StatusPatchFailureStopsScaleUp(t *testing.T) { + dep := newTestDeployment("openfga", "default", "openfga/openfga:v1.14.0", 0) + dep.Annotations[AnnotationDesiredReplicas] = "3" + dep.Status.Conditions = []appsv1.DeploymentCondition{{ + Type: "MigrationFailed", + Status: corev1.ConditionTrue, + }} + cm := &corev1.ConfigMap{ + ObjectMeta: metav1.ObjectMeta{ + Name: "openfga-migration-status", + Namespace: "default", + }, + Data: map[string]string{"version": "v1.14.0"}, + } + r := newReconcilerWithStatusPatchError(dep, cm) + + if _, err := r.Reconcile(context.Background(), ctrl.Request{ + NamespacedName: types.NamespacedName{Name: "openfga", Namespace: "default"}, + }); err == nil { + t.Fatal("expected status patch error") + } + + updated := &appsv1.Deployment{} + if err := r.Get(context.Background(), types.NamespacedName{ + Name: "openfga", Namespace: "default", + }, updated); err != nil { + t.Fatalf("getting deployment: %v", err) + } + if *updated.Spec.Replicas != 0 { + t.Errorf("expected replicas to remain at 0, got %d", *updated.Spec.Replicas) + } + cond := findCondition(updated.Status.Conditions, "MigrationFailed") + if cond == nil || cond.Status != corev1.ConditionTrue { + t.Fatalf("expected MigrationFailed condition to remain True, got %+v", cond) + } +} + func TestReconcile_JobSucceeded_UpdatesConfigMapAndScalesUp(t *testing.T) { // Given: a Deployment at 0 replicas, no ConfigMap, a succeeded migration Job, // and a pre-existing MigrationFailed condition from a prior attempt. @@ -372,6 +439,49 @@ func TestReconcile_JobFailed_SetsRetryAnnotationAndRequeues(t *testing.T) { } } +func TestReconcile_JobFailed_StatusPatchFailurePreservesJob(t *testing.T) { + dep := newTestDeployment("openfga", "default", "openfga/openfga:v1.14.0", 0) + dep.Annotations[AnnotationDesiredReplicas] = "3" + job := &batchv1.Job{ + ObjectMeta: metav1.ObjectMeta{ + Name: "openfga-migrate", + Namespace: "default", + Annotations: map[string]string{ + AnnotationDesiredVersion: "v1.14.0", + }, + }, + Status: batchv1.JobStatus{ + Conditions: []batchv1.JobCondition{{ + Type: batchv1.JobFailed, + Status: corev1.ConditionTrue, + }}, + }, + } + r := newReconcilerWithStatusPatchError(dep, job) + + if _, err := r.Reconcile(context.Background(), ctrl.Request{ + NamespacedName: types.NamespacedName{Name: "openfga", Namespace: "default"}, + }); err == nil { + t.Fatal("expected status patch error") + } + + preservedJob := &batchv1.Job{} + if err := r.Get(context.Background(), types.NamespacedName{ + Name: "openfga-migrate", Namespace: "default", + }, preservedJob); err != nil { + t.Fatalf("expected failed Job to remain for retry: %v", err) + } + updated := &appsv1.Deployment{} + if err := r.Get(context.Background(), types.NamespacedName{ + Name: "openfga", Namespace: "default", + }, updated); err != nil { + t.Fatalf("getting deployment: %v", err) + } + if _, ok := updated.Annotations[AnnotationRetryAfter]; ok { + t.Error("retry-after must not be set when the failure condition was not persisted") + } +} + func TestReconcile_JobFailureTarget_TreatedAsFailed(t *testing.T) { // Given: a Job with only JobFailureTarget=True (no JobFailed yet). The // Job controller sets this as soon as it decides the Job will fail, @@ -901,8 +1011,8 @@ func TestReconcile_JobSucceeded_UpdatesExistingConfigMap(t *testing.T) { Name: "openfga-migration-status", Namespace: "default", Labels: map[string]string{ - LabelPartOf: LabelPartOfValue, - LabelComponent: "migration", + LabelPartOf: LabelPartOfValue, + LabelComponent: "migration", "app.kubernetes.io/managed-by": "openfga-operator", }, OwnerReferences: []metav1.OwnerReference{ diff --git a/operator/tests/README.md b/operator/tests/README.md index e4377a99..e8426987 100644 --- a/operator/tests/README.md +++ b/operator/tests/README.md @@ -40,7 +40,7 @@ helm install openfga-test charts/openfga -n openfga-test \ |----------|-------| | `openfga-test-openfga-operator` | `1/1 Running` | | `openfga-test-postgres` | `1/1 Running` | -| `openfga-test-migrate-xxxxx` | `0/1 Completed` | +| `openfga-test-migrate` | `0/1 Completed` | | `openfga-test` (OpenFGA) | `3/3 Running` | **Verify:** From 317ff91a85db3a48c6cb841f057292f3e8854e43 Mon Sep 17 00:00:00 2001 From: Anurag Bandyopadhyay <angbpy@gmail.com> Date: Wed, 23 Sep 2026 09:58:32 +0530 Subject: [PATCH 47/70] operator: inherit pull policy, drop blockOwnerDeletion, strategic-merge status Migration Job inherits the OpenFGA container's imagePullPolicy, so it can't run a stale cached image while the app pulls a fresh one. Owner references no longer set blockOwnerDeletion. That requires the deployments/finalizers subresource the operator isn't granted, and the Job create fails under OwnerReferencesPermissionEnforcement. Controller: true still enables garbage collection. Status condition patches use StrategicMergeFrom, merging MigrationFailed by condition type instead of replacing the whole conditions list and clobbering Available/Progressing. --- operator/internal/controller/helpers.go | 23 +++++----- .../controller/migration_controller.go | 10 +++-- .../controller/migration_controller_test.go | 45 ++++++++++++++++++- 3 files changed, 60 insertions(+), 18 deletions(-) diff --git a/operator/internal/controller/helpers.go b/operator/internal/controller/helpers.go index d5852333..1c4f240c 100644 --- a/operator/internal/controller/helpers.go +++ b/operator/internal/controller/helpers.go @@ -155,12 +155,11 @@ func buildMigrationJob( }, OwnerReferences: []metav1.OwnerReference{ { - APIVersion: "apps/v1", - Kind: "Deployment", - Name: deployment.Name, - UID: deployment.UID, - Controller: ptr.To(true), - BlockOwnerDeletion: ptr.To(true), + APIVersion: "apps/v1", + Kind: "Deployment", + Name: deployment.Name, + UID: deployment.UID, + Controller: ptr.To(true), }, }, }, @@ -184,6 +183,7 @@ func buildMigrationJob( { Name: "migrate-database", Image: mainContainer.Image, + ImagePullPolicy: mainContainer.ImagePullPolicy, Args: []string{"migrate"}, Env: mainContainer.Env, EnvFrom: mainContainer.EnvFrom, @@ -223,12 +223,11 @@ func updateMigrationStatus( }, OwnerReferences: []metav1.OwnerReference{ { - APIVersion: "apps/v1", - Kind: "Deployment", - Name: deployment.Name, - UID: deployment.UID, - Controller: ptr.To(true), - BlockOwnerDeletion: ptr.To(true), + APIVersion: "apps/v1", + Kind: "Deployment", + Name: deployment.Name, + UID: deployment.UID, + Controller: ptr.To(true), }, }, }, diff --git a/operator/internal/controller/migration_controller.go b/operator/internal/controller/migration_controller.go index 177ce0e7..b89f3117 100644 --- a/operator/internal/controller/migration_controller.go +++ b/operator/internal/controller/migration_controller.go @@ -93,7 +93,9 @@ func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) ( // 5. If versions match, ensure Deployment is scaled up and return. if currentVersion == desiredVersion { logger.V(1).Info("migration up to date", "version", desiredVersion) - statusPatch := client.MergeFrom(deployment.DeepCopy()) + // Strategic merge patches the condition by type without replacing the whole + // conditions list, so it won't clobber Available/Progressing. + statusPatch := client.StrategicMergeFrom(deployment.DeepCopy()) if clearMigrationFailedCondition(deployment) { if patchErr := r.Status().Patch(ctx, deployment, statusPatch); patchErr != nil { logger.Error(patchErr, "failed to clear MigrationFailed condition") @@ -187,7 +189,7 @@ func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) ( logger.Info("migration succeeded", "version", desiredVersion) // Clear MigrationFailed condition. - statusPatch := client.MergeFrom(deployment.DeepCopy()) + statusPatch := client.StrategicMergeFrom(deployment.DeepCopy()) if clearMigrationFailedCondition(deployment) { if patchErr := r.Status().Patch(ctx, deployment, statusPatch); patchErr != nil { logger.Error(patchErr, "failed to clear MigrationFailed condition") @@ -215,8 +217,8 @@ func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) ( if isJobConditionTrue(job, batchv1.JobFailed) || isJobConditionTrue(job, batchv1.JobFailureTarget) { logger.Info("migration job failed, will delete and retry", "job", jobName, "version", desiredVersion) - // Set condition so kubectl describe shows the failure. - statusPatch := client.MergeFrom(deployment.DeepCopy()) + // Set MigrationFailed so kubectl describe shows the failure. + statusPatch := client.StrategicMergeFrom(deployment.DeepCopy()) if setMigrationFailedCondition(deployment, desiredVersion) { if patchErr := r.Status().Patch(ctx, deployment, statusPatch); patchErr != nil { logger.Error(patchErr, "failed to set MigrationFailed condition") diff --git a/operator/internal/controller/migration_controller_test.go b/operator/internal/controller/migration_controller_test.go index b7ab5317..5b921dec 100644 --- a/operator/internal/controller/migration_controller_test.go +++ b/operator/internal/controller/migration_controller_test.go @@ -138,6 +138,47 @@ func TestReconcile_FirstInstall_CreatesJob(t *testing.T) { } } +func TestReconcile_FirstInstall_JobInheritsPullPolicyAndOwnerRef(t *testing.T) { + // Given: a Deployment whose OpenFGA container pins an explicit pull policy. + dep := newTestDeployment("openfga", "default", "openfga/openfga:v1.14.0", 0) + dep.Spec.Template.Spec.Containers[0].ImagePullPolicy = corev1.PullAlways + r := newReconciler(dep) + + // When: reconciling. + if _, err := r.Reconcile(context.Background(), ctrl.Request{ + NamespacedName: types.NamespacedName{Name: "openfga", Namespace: "default"}, + }); err != nil { + t.Fatalf("unexpected error: %v", err) + } + + job := &batchv1.Job{} + if err := r.Get(context.Background(), types.NamespacedName{ + Name: "openfga-migrate", Namespace: "default", + }, job); err != nil { + t.Fatalf("expected migration job to be created: %v", err) + } + + // The migration Job must inherit the OpenFGA container's pull policy so it + // cannot run a stale cached image while the app pulls a fresh one. + if got := job.Spec.Template.Spec.Containers[0].ImagePullPolicy; got != corev1.PullAlways { + t.Errorf("expected job pull policy %q, got %q", corev1.PullAlways, got) + } + + // Owner reference makes the Job GC with the Deployment, but no blockOwnerDeletion: + // that needs the deployments/finalizers subresource the operator isn't granted, + // so the create would fail under OwnerReferencesPermissionEnforcement. + if len(job.OwnerReferences) != 1 { + t.Fatalf("expected exactly one owner reference, got %d", len(job.OwnerReferences)) + } + ref := job.OwnerReferences[0] + if ref.Controller == nil || !*ref.Controller { + t.Errorf("expected controller owner reference, got %+v", ref) + } + if ref.BlockOwnerDeletion != nil { + t.Errorf("expected blockOwnerDeletion to be unset, got %v", *ref.BlockOwnerDeletion) + } +} + func TestReconcile_VersionMatch_ScalesUp(t *testing.T) { // Given: a Deployment at 0 replicas with matching migration-status ConfigMap. dep := newTestDeployment("openfga", "default", "openfga/openfga:v1.14.0", 0) @@ -901,8 +942,8 @@ func TestReconcile_JobSucceeded_UpdatesExistingConfigMap(t *testing.T) { Name: "openfga-migration-status", Namespace: "default", Labels: map[string]string{ - LabelPartOf: LabelPartOfValue, - LabelComponent: "migration", + LabelPartOf: LabelPartOfValue, + LabelComponent: "migration", "app.kubernetes.io/managed-by": "openfga-operator", }, OwnerReferences: []metav1.OwnerReference{ From c11869cdb865bac00add337c1f94d3a4a73965ea Mon Sep 17 00:00:00 2001 From: Anurag Bandyopadhyay <angbpy@gmail.com> Date: Wed, 23 Sep 2026 09:58:40 +0530 Subject: [PATCH 48/70] chart: omit spec.replicas in operator mode, gate annotations on applyMigrations Operator mode omits spec.replicas rather than rendering a count. The operator owns it (scales up after the migration Job completes) and GitOps controllers treat an absent field as unmanaged, so neither fights the other and no ignoreDifferences is needed. A fresh install starts at the Kubernetes default of one replica, held NotReady by OpenFGA's readiness gate until migration. Operator annotations now also require datastore.applyMigrations: disabling it opts out of operator management and renders replicas normally. --- charts/openfga/templates/deployment.yaml | 15 ++++------ charts/openfga/tests/operator_mode_test.yaml | 30 ++++++++++++++------ 2 files changed, 26 insertions(+), 19 deletions(-) diff --git a/charts/openfga/templates/deployment.yaml b/charts/openfga/templates/deployment.yaml index 7d27a600..cebb05ec 100644 --- a/charts/openfga/templates/deployment.yaml +++ b/charts/openfga/templates/deployment.yaml @@ -4,7 +4,7 @@ metadata: name: {{ include "openfga.fullname" . }} labels: {{- include "openfga.labels" . | nindent 4 }} - {{- $hasOperatorAnnotations := and .Values.operator.enabled .Values.migration.enabled }} + {{- $hasOperatorAnnotations := and .Values.operator.enabled .Values.migration.enabled .Values.datastore.applyMigrations }} {{- if or $hasOperatorAnnotations .Values.annotations }} annotations: {{- if $hasOperatorAnnotations }} @@ -20,18 +20,13 @@ metadata: {{- end }} {{- end }} spec: - {{- if and .Values.operator.enabled .Values.migration.enabled }} + {{- if $hasOperatorAnnotations }} {{- if .Values.autoscaling.enabled }} {{- fail "operator.enabled and autoscaling.enabled cannot both be true" }} {{- end }} - {{- /* On upgrade: preserve live replicas (zero-downtime). On fresh install: lookup returns empty, fall back to 0. - OpenFGA gates readiness on MinimumSupportedDatastoreSchemaRevision — see ADR-002. */ -}} - {{- $existing := (lookup "apps/v1" "Deployment" (include "openfga.namespace" .) (include "openfga.fullname" .)) }} - {{- if and $existing (hasKey ($existing) "spec") }} - replicas: {{ $existing.spec.replicas }} - {{- else }} - replicas: {{ ternary 1 0 (eq .Values.datastore.engine "memory") }} - {{- end }} + {{- /* Operator mode: omit spec.replicas. The operator owns it (scales up after + the migration Job completes) and GitOps controllers treat an absent field + as unmanaged, so neither fights the other. See ADR-002. */ -}} {{- else if not .Values.autoscaling.enabled }} replicas: {{ ternary 1 .Values.replicaCount (eq .Values.datastore.engine "memory") }} {{- end }} diff --git a/charts/openfga/tests/operator_mode_test.yaml b/charts/openfga/tests/operator_mode_test.yaml index 164d8bd1..bdffcf03 100644 --- a/charts/openfga/tests/operator_mode_test.yaml +++ b/charts/openfga/tests/operator_mode_test.yaml @@ -64,30 +64,27 @@ tests: path: metadata.annotations["openfga.dev/migration-service-account"] # --- Replica count --- - # When no live cluster is available (helm template / test), lookup returns empty, - # so the template falls back to replicas: 0 (fresh install behavior). - # On a real cluster, lookup preserves the existing replica count for zero-downtime upgrades. - - it: should set replicas to 0 on fresh install when operator is enabled with database datastore + # Operator mode omits spec.replicas so the operator owns it and GitOps treats it + # as unmanaged. See ADR-002. + - it: should omit replicas in operator mode with a database datastore set: operator.enabled: true migration.enabled: true replicaCount: 3 datastore.engine: postgres asserts: - - equal: + - isNull: path: spec.replicas - value: 0 - - it: should set replicas to 1 when operator is enabled with memory datastore + - it: should omit replicas in operator mode with a memory datastore set: operator.enabled: true migration.enabled: true replicaCount: 5 datastore.engine: memory asserts: - - equal: + - isNull: path: spec.replicas - value: 1 - it: should set replicas to replicaCount when operator is disabled set: @@ -99,6 +96,21 @@ tests: path: spec.replicas value: 5 + # applyMigrations=false opts out: no operator annotations, replicas rendered normally. + - it: should not use operator mode when applyMigrations is false + set: + operator.enabled: true + migration.enabled: true + replicaCount: 4 + datastore.engine: postgres + datastore.applyMigrations: false + asserts: + - isNull: + path: metadata.annotations["openfga.dev/migration-enabled"] + - equal: + path: spec.replicas + value: 4 + # --- Autoscaling conflict --- - it: should fail when operator and autoscaling are both enabled set: From 5b42197c5d300021b4f76acc71cdde1b31b5dfde Mon Sep 17 00:00:00 2001 From: Anurag Bandyopadhyay <angbpy@gmail.com> Date: Wed, 23 Sep 2026 09:58:46 +0530 Subject: [PATCH 49/70] docs: operator migration replica model and limitations Update README and ADR-002 to describe omitting spec.replicas in operator mode instead of rendering 0, and how the readiness gate keeps a pre-migration pod from serving. Document that migrations key only on the image tag, the Job runs a single container (no sidecar-proxy support), and that the operator image tag defaults to the floating appVersion so it should be pinned in production. --- charts/openfga-operator/values.yaml | 4 ++- docs/adr/002-operator-managed-migrations.md | 27 ++++++++++++--------- operator/README.md | 26 +++++++------------- 3 files changed, 28 insertions(+), 29 deletions(-) diff --git a/charts/openfga-operator/values.yaml b/charts/openfga-operator/values.yaml index 921dcef7..a3375229 100644 --- a/charts/openfga-operator/values.yaml +++ b/charts/openfga-operator/values.yaml @@ -3,7 +3,9 @@ replicaCount: 1 image: repository: ghcr.io/openfga/openfga-operator pullPolicy: IfNotPresent - # -- Overrides the image tag whose default is the chart appVersion. + # -- Overrides the image tag (defaults to the chart appVersion). appVersion is a + # floating tag, so with pullPolicy IfNotPresent nodes can run different builds + # under it. Pin an immutable tag or digest in production. tag: "" imagePullSecrets: [] diff --git a/docs/adr/002-operator-managed-migrations.md b/docs/adr/002-operator-managed-migrations.md index 92a9abab..dd45e65f 100644 --- a/docs/adr/002-operator-managed-migrations.md +++ b/docs/adr/002-operator-managed-migrations.md @@ -88,23 +88,25 @@ The operator runs a **migration controller** that reconciles the OpenFGA Deploym │ └── ttlSecondsAfterFinished: 300 │ │ 5. Watch Job until succeeded │ │ 6. Update ConfigMap → "version: v1.14.0" │ -│ 7. Ensure Deployment at desired replicas │ -│ (fresh install: 0 → N; upgrade: already running) │ +│ 7. Scale Deployment to desired replicas │ +│ (fresh install: default 1 → N; upgrade: unchanged) │ │ 8. New pods pass readiness, serve requests │ └──────────────────────────────────────────────────────────┘ ``` **Key design decisions within this approach:** -#### Zero-downtime upgrades via lookup and readiness gating +#### Zero-downtime upgrades via omitted replicas and readiness gating -On **fresh install**, the Helm chart renders the Deployment with `replicas: 0` (no existing Deployment found via `lookup`). The operator runs the migration Job and scales the Deployment to the desired replica count afterward. +In operator mode the chart **omits `spec.replicas` entirely** rather than rendering a fixed number. The operator owns the replica count: it scales the Deployment to `openfga.dev/desired-replicas` once the migration Job succeeds. Omitting the field (the same treatment an HPA-managed Deployment gets) means a GitOps controller sees no declared replica count and leaves it unmanaged, so it never fights the operator over the value — no `ignoreDifferences` or field-ownership patch is required. -On **upgrade**, the chart uses Helm's `lookup` function to read the current replica count from the live Deployment and preserves it. Kubernetes starts a rolling update with the new image. OpenFGA has a **built-in schema version gate**: on startup, each instance calls `IsReady()` which checks the database schema revision against `MinimumSupportedDatastoreSchemaRevision` (via goose). If the schema is behind, the gRPC health endpoint returns `NOT_SERVING`, the readiness probe fails, and Kubernetes does not route traffic to the pod. Old pods continue serving on the migrated schema (OpenFGA migrations are additive/backward-compatible — this is how the existing Helm hook flow has operated for years with rolling updates). Once the operator's migration Job completes, new pods pass readiness and the rolling update proceeds. +On **fresh install**, no `spec.replicas` is set, so Kubernetes applies its default of one replica. That pod starts before the migration has run and is held `NotReady` by OpenFGA's readiness gate (see below), so it serves no traffic. The operator runs the migration Job and then scales the Deployment to the desired replica count. -This matches the existing zero-downtime behavior of the non-operator chart. The previous approach (always starting at `replicas: 0`) introduced a full outage on every `helm upgrade` — even for config-only changes — which was a regression from the existing rolling update model. +On **upgrade**, the field is still absent from the rendered manifest, so Kubernetes preserves the live replica count and starts a rolling update with the new image. OpenFGA has a **built-in schema version gate**: on startup, each instance calls `IsReady()` which checks the database schema revision against `MinimumSupportedDatastoreSchemaRevision` (via goose). If the schema is behind, the gRPC health endpoint returns `NOT_SERVING`, the readiness probe fails, and Kubernetes does not route traffic to the pod. Old pods continue serving on the migrated schema (OpenFGA migrations are additive/backward-compatible — this is how the existing Helm hook flow has operated for years with rolling updates). Once the operator's migration Job completes, new pods pass readiness and the rolling update proceeds. -**`lookup` caveat:** `helm template` and `--dry-run=client` cannot query the cluster, so `lookup` returns empty and the template falls back to `replicas: 0`. This is correct for CI rendering (no live cluster) and does not affect real installs/upgrades. `--dry-run=server` works correctly. +This matches the existing zero-downtime behavior of the non-operator chart. + +**Rejected alternative — pin `replicas: 0` and read the live count via `lookup`:** the chart could render `replicas: 0` on fresh install and use Helm's `lookup` to preserve the live count on upgrade. This is rejected for two reasons. A GitOps controller that reconciles the manifest drives replicas back to 0 on every sync and fights the operator, causing a full outage. And `lookup` returns empty under `helm template` and `--dry-run=client`, so the manifest silently renders `replicas: 0` in CI and dry runs. Omitting the field avoids both problems and needs no cluster lookup. #### Version tracking via ConfigMap @@ -155,10 +157,13 @@ Problems: ArgoCD skips step 4. FluxCD deletes Job in step 4. `--wait` deadlocks helm install ├── Create ServiceAccount (runtime), ServiceAccount (migrator) ├── Create Secret, Service - ├── Create Deployment (replicas: 0 via lookup fallback, no init containers) + ├── Create Deployment (no spec.replicas, no init containers) ├── Create Operator Deployment └── [Helm is done — all resources are regular, no hooks] +(Kubernetes starts the Deployment at its default of 1 replica; that pod +is held NotReady by the readiness gate until the migration completes.) + Operator starts: ├── Detects Deployment image version ├── No migration status ConfigMap → migration needed @@ -166,14 +171,14 @@ Operator starts: │ └── Uses openfga-migrator ServiceAccount │ └── Runs openfga migrate → succeeds ├── Creates ConfigMap with migrated version - └── Scales Deployment 0 → 3 replicas → pods start + └── Scales Deployment to 3 replicas → pods pass readiness ``` **After (operator-managed, upgrade with new image):** ``` helm upgrade - ├── lookup finds existing Deployment at 3 replicas → preserves replicas: 3 + ├── no spec.replicas in manifest → Kubernetes keeps the live count (3) ├── Patches Deployment with new image tag ├── Kubernetes starts rolling update │ ├── New pods (v1.14) start → schema is behind → @@ -232,7 +237,7 @@ Users on `operator.enabled: false` (the default) see identical rendered output t ### Negative - **Operator is a new runtime dependency** — if the operator pod is unavailable, migrations don't run (but existing running pods are unaffected) -- **`lookup` limitation** — `helm template` and `--dry-run=client` cannot query the cluster; the template falls back to `replicas: 0` in these contexts. This does not affect real installs/upgrades. +- **Replica count is unmanaged by the chart in operator mode** — because `spec.replicas` is omitted, the rendered manifest no longer declares a desired count; the operator (and, on first install, the Kubernetes default of 1) determines it. A reader inspecting only the chart output cannot see the running replica count. - **Two upgrade paths to document** — `operator.enabled: true` (new) vs `operator.enabled: false` (legacy) ### Risks diff --git a/operator/README.md b/operator/README.md index ea03c587..b0fc515e 100644 --- a/operator/README.md +++ b/operator/README.md @@ -8,12 +8,13 @@ This is **Stage 1** of the operator — focused solely on migration orchestratio 1. The operator watches Deployments **in its own namespace** labeled `app.kubernetes.io/part-of: openfga` and `app.kubernetes.io/component: authorization-controller` 2. When a version change is detected (comparing the container image tag to the `{name}-migration-status` ConfigMap), the operator: - - Keeps the Deployment at 0 replicas - Creates a migration Job running `openfga migrate` - Waits for the Job to complete - Updates the ConfigMap with the new version - - Scales the Deployment up to the desired replica count -3. On failure, a `MigrationFailed` condition is set on the Deployment and replicas stay at 0 + - Scales the Deployment to the desired replica count (`openfga.dev/desired-replicas`) +3. On failure, a `MigrationFailed` condition is set on the Deployment and the desired replica count is not applied + +The operator never scales the Deployment to 0. A pod that starts before the migration completes is held `NotReady` by OpenFGA's readiness gate on `MinimumSupportedDatastoreSchemaRevision`, so it won't serve traffic against an unmigrated schema. ## Prerequisites @@ -124,23 +125,14 @@ The operator reads these annotations from the OpenFGA Deployment: | Annotation | Description | |------------|-------------| -| `openfga.dev/migration-enabled` | Must be `"true"` for the operator to manage migrations. Deployments without this annotation are ignored. Set by the Helm chart when `operator.enabled` and `migration.enabled` are both true. | -| `openfga.dev/desired-replicas` | The replica count to restore after migration succeeds. Set by the Helm chart. | +| `openfga.dev/migration-enabled` | Must be `"true"` for the operator to manage migrations. Deployments without this annotation are ignored. Set by the Helm chart when `operator.enabled`, `migration.enabled`, and `datastore.applyMigrations` are all true. | +| `openfga.dev/desired-replicas` | The replica count the operator scales the Deployment to once migration succeeds. Set by the Helm chart. | | `openfga.dev/migration-service-account` | The ServiceAccount to use for migration Jobs. Defaults to the Deployment's SA. | ## Limitations -- **Mutable image tags:** The operator detects version changes by comparing the container image tag (or digest). If you deploy with a mutable tag like `latest` or reuse the same tag for different builds, the operator will not detect changes and will skip the migration. Use immutable tags (e.g., `v1.14.0`) or pin images by digest for reliable migration triggering. +- **Migrations key only on the image tag:** The operator compares the container image tag (or digest) to the `{name}-migration-status` ConfigMap. A mutable tag like `latest`, or a tag reused for a new build, is not seen as a change, so the migration is skipped — use immutable tags (e.g. `v1.14.0`) or pin by digest. A migration-needing change that keeps the same image — for example repointing `datastore.uri` at a different or restored database — also won't trigger a Job; the readiness gate holds the new pod `NotReady`, but you must migrate manually (bump the image or delete the status ConfigMap). - **Migration-specific volumes:** The legacy Helm chart values `migrate.extraVolumes` and `migrate.extraVolumeMounts` have no effect in operator mode. The operator inherits volumes and mounts from the main Deployment pod spec. If you need additional volumes for migrations (e.g., CA bundles or TLS certs), add them to the top-level `extraVolumes` and `extraVolumeMounts` values instead. +- **Single-container migration Job:** The Job runs one container (`openfga migrate`) with the main container's env, volumes, and scheduling. It injects no sidecars or extra init containers, so databases reached through a sidecar proxy (Cloud SQL Auth Proxy, AlloyDB) aren't supported for operator-managed migrations — a proxy that doesn't exit on its own (e.g. an Istio sidecar) would keep the Job pod running and stop the Job from completing. Connect to such databases directly instead. - **`envFrom` datastore detection:** The memory-datastore check inspects only the explicit `env` entries on the container. If `OPENFGA_DATASTORE_ENGINE` is supplied via `envFrom` (a ConfigMap or Secret), the operator cannot read the value and will attempt a migration Job that a memory datastore does not need. The Helm chart sets this variable inline, so chart-managed installs are unaffected. -- **GitOps and `spec.replicas`:** The operator owns the Deployment's replica count — it scales to `0` for the duration of a migration and restores `openfga.dev/desired-replicas` afterward. Under a GitOps controller the chart's `lookup` of the live replica count returns empty at render time, so the rendered manifest carries `replicas: 0` (see [ADR-002](../docs/adr/002-operator-managed-migrations.md)). If the controller keeps syncing that field it will fight the operator over it, so exclude `spec.replicas` from GitOps reconciliation — the same treatment an HPA-managed Deployment needs: - - **ArgoCD** — add to the `Application`: - ```yaml - spec: - ignoreDifferences: - - group: apps - kind: Deployment - jsonPointers: - - /spec/replicas - ``` - - **FluxCD** — Flux applies server-side and honors field ownership, so drop `spec.replicas` from the Flux-applied manifest (e.g. a Kustomize patch removing `/spec/replicas`) and let the operator own it. +- **GitOps and `spec.replicas`:** The operator owns the replica count — it scales to `openfga.dev/desired-replicas` after the migration Job completes (see [ADR-002](../docs/adr/002-operator-managed-migrations.md)). To avoid fighting a GitOps controller over that field, the chart omits `spec.replicas` in operator mode, the same as an HPA-managed Deployment. Argo CD and Flux treat an absent field as unmanaged, so no `ignoreDifferences` or field-ownership patch is needed. A fresh install starts at the Kubernetes default of one replica, held `NotReady` by the readiness gate until the migration finishes, after which the operator scales to the desired count. From 504a9793406b5e079c76fe52fb6a5078526b0095 Mon Sep 17 00:00:00 2001 From: Anurag Bandyopadhyay <angbpy@gmail.com> Date: Wed, 23 Sep 2026 11:36:45 +0530 Subject: [PATCH 50/70] operator CI: publish the version image tag once per version build-and-push re-published :<version> on every push to main, so the tag the chart deploys by default could resolve to different builds over time. Publish :<version> only when it is not already in the registry (skip-existing, matching chart-releaser's CR_SKIP_EXISTING for the charts); keep refreshing :latest and always publish the immutable :<version>-<sha>. The chart's default image tag is now stable for a given version. --- .github/workflows/operator.yml | 54 ++++++++++++++++++++++------- charts/openfga-operator/values.yaml | 8 ++--- 2 files changed, 45 insertions(+), 17 deletions(-) diff --git a/.github/workflows/operator.yml b/.github/workflows/operator.yml index a09638c9..ed49c5f5 100644 --- a/.github/workflows/operator.yml +++ b/.github/workflows/operator.yml @@ -69,24 +69,18 @@ jobs: echo "short_sha=${short_sha}" >> "$GITHUB_OUTPUT" echo "Operator version: ${version} (sha: ${short_sha})" - - name: Determine image tags and push policy - id: tags + - name: Determine push policy + id: policy run: | if [[ "${{ github.event_name }}" == "push" && "${{ github.ref }}" == "refs/heads/main" ]]; then - # Main push: publish floating :<version> and :latest plus an - # immutable :<version>-<sha> so consumers pinning a specific - # commit have a stable reference. - echo "tags=${{ env.IMAGE_NAME }}:${{ steps.version.outputs.version }},${{ env.IMAGE_NAME }}:latest,${{ env.IMAGE_NAME }}:${{ steps.version.outputs.version }}-${{ steps.version.outputs.short_sha }}" >> "$GITHUB_OUTPUT" echo "push=true" >> "$GITHUB_OUTPUT" + echo "mode=main" >> "$GITHUB_OUTPUT" elif [[ "${{ github.event_name }}" == "workflow_dispatch" && "${{ inputs.push_image }}" == "true" ]]; then - echo "tags=${{ env.IMAGE_NAME }}:${{ steps.version.outputs.version }}-${{ steps.version.outputs.short_sha }}" >> "$GITHUB_OUTPUT" echo "push=true" >> "$GITHUB_OUTPUT" + echo "mode=dispatch" >> "$GITHUB_OUTPUT" else - # Pull request (or workflow_dispatch with push_image=false): - # build both platforms but don't publish — catches arm64-incompatible - # changes (build tags, syscalls, CGO) before they merge. - echo "tags=${{ env.IMAGE_NAME }}:pr-${{ steps.version.outputs.short_sha }}" >> "$GITHUB_OUTPUT" echo "push=false" >> "$GITHUB_OUTPUT" + echo "mode=pr" >> "$GITHUB_OUTPUT" fi - name: Set up QEMU @@ -96,18 +90,52 @@ jobs: uses: docker/setup-buildx-action@v3 - name: Login to GHCR - if: steps.tags.outputs.push == 'true' + if: steps.policy.outputs.push == 'true' uses: docker/login-action@v4.1.0 with: registry: ghcr.io username: ${{ github.actor }} password: ${{ secrets.GITHUB_TOKEN }} + - name: Determine image tags + id: tags + run: | + set -euo pipefail + version="${{ steps.version.outputs.version }}" + sha="${{ steps.version.outputs.short_sha }}" + img="${{ env.IMAGE_NAME }}" + case "${{ steps.policy.outputs.mode }}" in + main) + # Always refresh :latest and publish the immutable per-commit tag. + tags="${img}:latest,${img}:${version}-${sha}" + # Publish :<version> only if it is not already in the registry, so a + # released version is never overwritten by a later push to main. This + # mirrors chart-releaser's CR_SKIP_EXISTING for the charts and keeps + # the chart's default image tag stable for a given version. + if docker buildx imagetools inspect "${img}:${version}" >/dev/null 2>&1; then + echo "::notice::${img}:${version} already exists; leaving it unchanged" + else + tags="${img}:${version},${tags}" + fi + ;; + dispatch) + # Manual run: publish only the immutable per-commit tag. + tags="${img}:${version}-${sha}" + ;; + *) + # Pull request: build both platforms but do not publish — catches + # arm64-incompatible changes (build tags, syscalls, CGO) before merge. + tags="${img}:pr-${sha}" + ;; + esac + echo "tags=${tags}" >> "$GITHUB_OUTPUT" + echo "Resolved tags: ${tags}" + - name: Build and (conditionally) push uses: docker/build-push-action@v6 with: context: operator - push: ${{ steps.tags.outputs.push }} + push: ${{ steps.policy.outputs.push }} platforms: linux/amd64,linux/arm64 tags: ${{ steps.tags.outputs.tags }} cache-from: type=gha diff --git a/charts/openfga-operator/values.yaml b/charts/openfga-operator/values.yaml index 08de2645..8e29275c 100644 --- a/charts/openfga-operator/values.yaml +++ b/charts/openfga-operator/values.yaml @@ -3,10 +3,10 @@ replicaCount: 1 image: repository: ghcr.io/openfga/openfga-operator pullPolicy: IfNotPresent - # -- Overrides the image tag (defaults to the chart appVersion). The appVersion - # tag is floating — CI republishes it on every push to main — so with pullPolicy - # IfNotPresent nodes can end up on different builds under the same tag. For - # production, pin the immutable `<appVersion>-<git-sha>` tag or a digest. + # -- Overrides the image tag (defaults to the chart appVersion). CI publishes the + # appVersion tag once per version and does not overwrite it, so it stays stable + # for a given chart version. To pin a specific main build instead, use the + # immutable `<appVersion>-<git-sha>` tag or a digest. tag: "" imagePullSecrets: [] From d46096c11fbaf0fe70a7ab3ef3fad732f17ebd20 Mon Sep 17 00:00:00 2001 From: SoulPancake <angbpy@gmail.com> Date: Wed, 23 Sep 2026 14:04:06 +0530 Subject: [PATCH 51/70] chart: keep legacy migration init containers identical to main The operator guard also required applyMigrations and waitForMigrations for the initContainer migration mode. With migrationType=initContainer and extraInitContainers set, main still renders the migrate-database init container when waitForMigrations is false, so upgrading the chart silently stopped those releases from migrating. Keep main's conditions and only add the operator.enabled check. --- charts/openfga/templates/deployment.yaml | 10 ++++------ 1 file changed, 4 insertions(+), 6 deletions(-) diff --git a/charts/openfga/templates/deployment.yaml b/charts/openfga/templates/deployment.yaml index cebb05ec..513eded8 100644 --- a/charts/openfga/templates/deployment.yaml +++ b/charts/openfga/templates/deployment.yaml @@ -55,11 +55,10 @@ spec: serviceAccountName: {{ include "openfga.serviceAccountName" . }} securityContext: {{- toYaml .Values.podSecurityContext | nindent 8 }} - {{- $needsMigrationInit := and (not .Values.operator.enabled) (has .Values.datastore.engine (list "postgres" "mysql")) .Values.datastore.applyMigrations .Values.datastore.waitForMigrations }} - {{- if or $needsMigrationInit .Values.extraInitContainers }} + {{- $legacyMigrations := and (not .Values.operator.enabled) (has .Values.datastore.engine (list "postgres" "mysql")) }} + {{ if or (and $legacyMigrations .Values.datastore.applyMigrations .Values.datastore.waitForMigrations) .Values.extraInitContainers }} initContainers: - {{- if $needsMigrationInit }} - {{- if eq .Values.datastore.migrationType "job" }} + {{- if and $legacyMigrations .Values.datastore.applyMigrations .Values.datastore.waitForMigrations (eq .Values.datastore.migrationType "job") }} - name: wait-for-migration securityContext: {{- toYaml .Values.securityContext | nindent 12 }} @@ -69,7 +68,7 @@ spec: resources: {{- toYaml .Values.datastore.migrations.resources | nindent 12 }} {{- end }} - {{- if eq .Values.datastore.migrationType "initContainer" }} + {{- if and $legacyMigrations (eq .Values.datastore.migrationType "initContainer") }} {{- with .Values.migrate.extraInitContainers }} {{- toYaml . | nindent 8 }} {{- end }} @@ -97,7 +96,6 @@ spec: {{- include "common.tplvalues.render" ( dict "value" .Values.migrate.sidecars "context" $) | nindent 8 }} {{- end }} {{- end }} - {{- end }} {{- with .Values.extraInitContainers }} {{- toYaml . | nindent 8 }} {{- end }} From 833a4a8c38921d00e709be310829679d449e6153 Mon Sep 17 00:00:00 2001 From: SoulPancake <angbpy@gmail.com> Date: Wed, 23 Sep 2026 14:04:47 +0530 Subject: [PATCH 52/70] operator chart: allow the operator to record events Leader election records a LeaderElection event on its Lease, and the API server rejected it ("events is forbidden") every time the operator took the lease. --- charts/openfga-operator/templates/role.yaml | 3 +++ 1 file changed, 3 insertions(+) diff --git a/charts/openfga-operator/templates/role.yaml b/charts/openfga-operator/templates/role.yaml index fa2863a7..eae4d873 100644 --- a/charts/openfga-operator/templates/role.yaml +++ b/charts/openfga-operator/templates/role.yaml @@ -21,3 +21,6 @@ rules: - apiGroups: ["coordination.k8s.io"] resources: ["leases"] verbs: ["get", "list", "watch", "create", "update"] + - apiGroups: [""] + resources: ["events"] + verbs: ["create", "patch"] From 5a6fcde2e2bd9c14872f40493e49f79b4e8ad363 Mon Sep 17 00:00:00 2001 From: SoulPancake <angbpy@gmail.com> Date: Wed, 23 Sep 2026 14:04:57 +0530 Subject: [PATCH 53/70] operator: build with Go 1.26.8 and bump golang.org/x/net and x/text govulncheck reported 12 vulnerabilities reachable from the operator: standard library issues in net/http, net/url, crypto/tls, crypto/x509 and encoding/asn1 fixed by Go 1.26.6, plus golang.org/x/net (fixed in v0.55.0) and golang.org/x/text (fixed in v0.39.0). It reports none after this change. --- operator/Dockerfile | 4 ++-- operator/README.md | 2 +- operator/go.mod | 12 ++++++------ operator/go.sum | 28 ++++++++++++++-------------- 4 files changed, 23 insertions(+), 23 deletions(-) diff --git a/operator/Dockerfile b/operator/Dockerfile index 034e22d0..1d1d71c4 100644 --- a/operator/Dockerfile +++ b/operator/Dockerfile @@ -1,5 +1,5 @@ -# pinned multi-arch index for golang:1.26.2 (linux/amd64, linux/arm64, ...) -FROM --platform=$BUILDPLATFORM golang:1.26.2@sha256:5f3787b7f902c07c7ec4f3aa91a301a3eda8133aa32661a3b3a3a86ab3a68a36 AS builder +# pinned multi-arch index for golang:1.26.8 (linux/amd64, linux/arm64, ...) +FROM --platform=$BUILDPLATFORM golang:1.26.8@sha256:6c2a5538f964f1c82f97ad14988bf05de100d922d159d0e398b54c7b0ca0c6c9 AS builder # buildx provides these automatically; declare so Go cross-compiles to the # requested target instead of the build host's arch. diff --git a/operator/README.md b/operator/README.md index fa4fc131..480b740b 100644 --- a/operator/README.md +++ b/operator/README.md @@ -18,7 +18,7 @@ The operator never scales the Deployment to 0. A pod that starts before the migr ## Prerequisites -- Go 1.26.2+ +- Go 1.26.8+ - Docker - Helm 3.6+ - A Kubernetes cluster (Rancher Desktop, kind, etc.) diff --git a/operator/go.mod b/operator/go.mod index 8cd2bbe4..72c27f7e 100644 --- a/operator/go.mod +++ b/operator/go.mod @@ -1,6 +1,6 @@ module github.com/openfga/openfga-operator -go 1.26.2 +go 1.26.8 require ( k8s.io/api v0.35.3 @@ -44,12 +44,12 @@ require ( go.uber.org/zap v1.27.0 // indirect go.yaml.in/yaml/v2 v2.4.3 // indirect go.yaml.in/yaml/v3 v3.0.4 // indirect - golang.org/x/net v0.47.0 // indirect + golang.org/x/net v0.59.0 // indirect golang.org/x/oauth2 v0.30.0 // indirect - golang.org/x/sync v0.18.0 // indirect - golang.org/x/sys v0.38.0 // indirect - golang.org/x/term v0.37.0 // indirect - golang.org/x/text v0.31.0 // indirect + golang.org/x/sync v0.23.0 // indirect + golang.org/x/sys v0.48.0 // indirect + golang.org/x/term v0.46.0 // indirect + golang.org/x/text v0.42.0 // indirect golang.org/x/time v0.9.0 // indirect gomodules.xyz/jsonpatch/v2 v2.4.0 // indirect google.golang.org/protobuf v1.36.8 // indirect diff --git a/operator/go.sum b/operator/go.sum index 79e74816..b321fc24 100644 --- a/operator/go.sum +++ b/operator/go.sum @@ -113,24 +113,24 @@ go.yaml.in/yaml/v2 v2.4.3 h1:6gvOSjQoTB3vt1l+CU+tSyi/HOjfOjRLJ4YwYZGwRO0= go.yaml.in/yaml/v2 v2.4.3/go.mod h1:zSxWcmIDjOzPXpjlTTbAsKokqkDNAVtZO0WOMiT90s8= go.yaml.in/yaml/v3 v3.0.4 h1:tfq32ie2Jv2UxXFdLJdh3jXuOzWiL1fo0bu/FbuKpbc= go.yaml.in/yaml/v3 v3.0.4/go.mod h1:DhzuOOF2ATzADvBadXxruRBLzYTpT36CKvDb3+aBEFg= -golang.org/x/mod v0.29.0 h1:HV8lRxZC4l2cr3Zq1LvtOsi/ThTgWnUk/y64QSs8GwA= -golang.org/x/mod v0.29.0/go.mod h1:NyhrlYXJ2H4eJiRy/WDBO6HMqZQ6q9nk4JzS3NuCK+w= -golang.org/x/net v0.47.0 h1:Mx+4dIFzqraBXUugkia1OOvlD6LemFo1ALMHjrXDOhY= -golang.org/x/net v0.47.0/go.mod h1:/jNxtkgq5yWUGYkaZGqo27cfGZ1c5Nen03aYrrKpVRU= +golang.org/x/mod v0.41.0 h1:qJmnOUb4YB+FsEuM3HcWucdZASCPGhsX6uljO6pog0c= +golang.org/x/mod v0.41.0/go.mod h1:Ek9pY8RKWXwsWvd3rQiHYtMqkjSUV+s1Rj7j4H5Ur6o= +golang.org/x/net v0.59.0 h1:5zfYln+w5XCxwrnMMJPufRgNoXEaGxl0wo5GqPXyues= +golang.org/x/net v0.59.0/go.mod h1:2DA/G1UfVbCpQPeWTmMPGY7Cs2PkBkwu743bVX5PIVg= golang.org/x/oauth2 v0.30.0 h1:dnDm7JmhM45NNpd8FDDeLhK6FwqbOf4MLCM9zb1BOHI= golang.org/x/oauth2 v0.30.0/go.mod h1:B++QgG3ZKulg6sRPGD/mqlHQs5rB3Ml9erfeDY7xKlU= -golang.org/x/sync v0.18.0 h1:kr88TuHDroi+UVf+0hZnirlk8o8T+4MrK6mr60WkH/I= -golang.org/x/sync v0.18.0/go.mod h1:9KTHXmSnoGruLpwFjVSX0lNNA75CykiMECbovNTZqGI= -golang.org/x/sys v0.38.0 h1:3yZWxaJjBmCWXqhN1qh02AkOnCQ1poK6oF+a7xWL6Gc= -golang.org/x/sys v0.38.0/go.mod h1:OgkHotnGiDImocRcuBABYBEXf8A9a87e/uXjp9XT3ks= -golang.org/x/term v0.37.0 h1:8EGAD0qCmHYZg6J17DvsMy9/wJ7/D/4pV/wfnld5lTU= -golang.org/x/term v0.37.0/go.mod h1:5pB4lxRNYYVZuTLmy8oR2BH8dflOR+IbTYFD8fi3254= -golang.org/x/text v0.31.0 h1:aC8ghyu4JhP8VojJ2lEHBnochRno1sgL6nEi9WGFGMM= -golang.org/x/text v0.31.0/go.mod h1:tKRAlv61yKIjGGHX/4tP1LTbc13YSec1pxVEWXzfoeM= +golang.org/x/sync v0.23.0 h1:KameEIfc1IkluZyXWLn39Wd4tURc6GbCiISGiZm2bQk= +golang.org/x/sync v0.23.0/go.mod h1:sUUOizhqBxiL6pEWpqNLUiaJn1ShEbZ6BBqskPbjZm0= +golang.org/x/sys v0.48.0 h1:bbX/i/6MgT9BVLM9RT1thmxL04yeTAhbEz4SyadbXoo= +golang.org/x/sys v0.48.0/go.mod h1:hNLxWAXmnKAxqDtdwIYC4bM9oQPEecfsnNMuSxOs3og= +golang.org/x/term v0.46.0 h1:3+OXuTbaKDgwk8jTi3aSLHRlmWqHEUDUtxnbFigO4YE= +golang.org/x/term v0.46.0/go.mod h1:+K02xbkittuwc0Am4abfA3Fc+XRGXkvBXNO88NCXPoc= +golang.org/x/text v0.42.0 h1:JbOZXgfeCPU9gacVtYliJqOhD+zhrEqK4LfdpmlUZqI= +golang.org/x/text v0.42.0/go.mod h1:ojzP1Z+2QtioaF8DTtO8K5q7JWVVYwZKenzujK0Zd0E= golang.org/x/time v0.9.0 h1:EsRrnYcQiGH+5FfbgvV4AP7qEZstoyrHB0DzarOQ4ZY= golang.org/x/time v0.9.0/go.mod h1:3BpzKBy/shNhVucY/MWOyx10tF3SFh9QdLuxbVysPQM= -golang.org/x/tools v0.38.0 h1:Hx2Xv8hISq8Lm16jvBZ2VQf+RLmbd7wVUsALibYI/IQ= -golang.org/x/tools v0.38.0/go.mod h1:yEsQ/d/YK8cjh0L6rZlY8tgtlKiBNTL14pGDJPJpYQs= +golang.org/x/tools v0.49.0 h1:3NI7VXzL9+1WZD52Dx2ttoPwD5DWrFGpl9mFZDlmisI= +golang.org/x/tools v0.49.0/go.mod h1:SJNXV9DBKT0UbdttsQjbfJlAE/q+y36++zo3uL3N0Oo= gomodules.xyz/jsonpatch/v2 v2.4.0 h1:Ci3iUJyx9UeRx7CeFN8ARgGbkESwJK+KB9lLcWxY/Zw= gomodules.xyz/jsonpatch/v2 v2.4.0/go.mod h1:AH3dM2RI6uoBZxn3LVrfvJ3E0/9dG4cSrbuBJT4moAY= google.golang.org/protobuf v1.36.8 h1:xHScyCOEuuwZEc6UtSOvPbAT4zRh0xcNRYekJwfqyMc= From 1659b951ccca14156aa1abcf2b044b13ed2b7112 Mon Sep 17 00:00:00 2001 From: SoulPancake <angbpy@gmail.com> Date: Wed, 23 Sep 2026 14:05:06 +0530 Subject: [PATCH 54/70] ci: pin operator workflow actions and add Dependabot for the operator Pin the actions in operator.yml to commit SHAs like the other workflows (#344), reusing the checkout and login-action versions already pinned there. Add gomod and docker entries for operator/ so the module and base images get updates. --- .github/dependabot.yaml | 18 ++++++++++++++++++ .github/workflows/operator.yml | 14 +++++++------- 2 files changed, 25 insertions(+), 7 deletions(-) diff --git a/.github/dependabot.yaml b/.github/dependabot.yaml index a561bab4..7560ce9e 100644 --- a/.github/dependabot.yaml +++ b/.github/dependabot.yaml @@ -19,3 +19,21 @@ updates: dependencies: patterns: - "*" + + - package-ecosystem: "gomod" + directory: "/operator" + schedule: + interval: "weekly" + groups: + dependencies: + patterns: + - "*" + + - package-ecosystem: "docker" + directory: "/operator" + schedule: + interval: "weekly" + groups: + dependencies: + patterns: + - "*" diff --git a/.github/workflows/operator.yml b/.github/workflows/operator.yml index ed49c5f5..6a2bc72d 100644 --- a/.github/workflows/operator.yml +++ b/.github/workflows/operator.yml @@ -30,10 +30,10 @@ jobs: contents: read steps: - name: Checkout - uses: actions/checkout@v6 + uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1 - name: Set up Go - uses: actions/setup-go@v5 + uses: actions/setup-go@40f1582b2485089dde7abd97c1529aa768e1baff # v5.6.0 with: go-version-file: operator/go.mod cache-dependency-path: operator/go.sum @@ -58,7 +58,7 @@ jobs: packages: write steps: - name: Checkout - uses: actions/checkout@v6 + uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1 - name: Extract version from Chart.yaml id: version @@ -84,14 +84,14 @@ jobs: fi - name: Set up QEMU - uses: docker/setup-qemu-action@v3 + uses: docker/setup-qemu-action@c7c53464625b32c7a7e944ae62b3e17d2b600130 # v3.7.0 - name: Set up Docker Buildx - uses: docker/setup-buildx-action@v3 + uses: docker/setup-buildx-action@8d2750c68a42422c14e847fe6c8ac0403b4cbd6f # v3.12.0 - name: Login to GHCR if: steps.policy.outputs.push == 'true' - uses: docker/login-action@v4.1.0 + uses: docker/login-action@dbcb813823bdd20940b903addbd779551569679f # v4.6.0 with: registry: ghcr.io username: ${{ github.actor }} @@ -132,7 +132,7 @@ jobs: echo "Resolved tags: ${tags}" - name: Build and (conditionally) push - uses: docker/build-push-action@v6 + uses: docker/build-push-action@10e90e3645eae34f1e60eeb005ba3a3d33f178e8 # v6.19.2 with: context: operator push: ${{ steps.policy.outputs.push }} From 8e7e566a99a1d384bd523f46107618ba05615667 Mon Sep 17 00:00:00 2001 From: SoulPancake <angbpy@gmail.com> Date: Wed, 23 Sep 2026 14:11:21 +0530 Subject: [PATCH 55/70] operator: only run migrations, leave the replica count to the chart The operator scaled the Deployment to an openfga.dev/desired-replicas annotation and the chart omitted spec.replicas, which broke in practice: - Switching an existing release to operator mode removed spec.replicas, so Helm's three-way merge and server-side apply both reset the Deployment to one replica until the migration finished. - kubectl scale was reverted immediately and HPAs could not be used. - The readiness gate it relied on only holds pods back below schema revision 4, so on upgrades new pods already served before the migration ran. The chart now renders replicas as in legacy mode and the operator only runs migration Jobs. Along with that: - Trust a migration Job only by its openfga.dev/desired-version annotation. The label fallback accepted the legacy Helm hook Job, whose version label is the chart appVersion, and skipped migration 006 on an upgrade from v1.9.5. - Keep a failed Job for 60s and then replace it, instead of tracking the delay in a retry-after annotation on the Deployment. - Leave activeDeadlineSeconds unset by default, as the hook Job does, so long index builds and table rebuilds are not killed halfway. A Job whose pod cannot start is rebuilt instead once the Deployment's pod template changes, tracked through an openfga.dev/pod-template-hash annotation. - Only opt the Deployment in, and create the migration service account, for Postgres and MySQL, which lets the operator drop its memory engine check. --- .github/ci/operator-postgres-values.yaml | 8 +- .github/workflows/test.yml | 4 +- charts/openfga-operator/templates/role.yaml | 2 +- charts/openfga-operator/values.schema.json | 2 +- charts/openfga-operator/values.yaml | 4 +- charts/openfga/ci/operator-mode-values.yaml | 18 +- charts/openfga/templates/NOTES.txt | 19 +- charts/openfga/templates/_helpers.tpl | 9 + charts/openfga/templates/deployment.yaml | 16 +- charts/openfga/templates/serviceaccount.yaml | 2 +- .../operator_mode_serviceaccount_test.yaml | 14 + charts/openfga/tests/operator_mode_test.yaml | 46 +- charts/openfga/values.yaml | 4 +- docs/adr/001-adopt-openfga-operator.md | 4 +- docs/adr/002-operator-managed-migrations.md | 56 +- operator/README.md | 28 +- operator/cmd/main.go | 2 +- operator/internal/controller/helpers.go | 276 ++-- .../controller/migration_controller.go | 338 ++-- .../controller/migration_controller_test.go | 1410 ++++------------- operator/tests/README.md | 25 +- 21 files changed, 641 insertions(+), 1646 deletions(-) diff --git a/.github/ci/operator-postgres-values.yaml b/.github/ci/operator-postgres-values.yaml index 4330d9ff..842566e0 100644 --- a/.github/ci/operator-postgres-values.yaml +++ b/.github/ci/operator-postgres-values.yaml @@ -1,8 +1,6 @@ -# E2E values consumed by the "operator + postgres E2E" step in test.yml. -# Not under charts/openfga/ci/ on purpose — chart-testing's helm-test runs -# a gRPC probe immediately after install, which would race the operator's -# scale-up. The dedicated workflow step waits for the migration ConfigMap -# and the scale-up explicitly, then verifies readiness. +# Values for the operator + Postgres E2E step in .github/workflows/test.yml, +# which installs one OpenFGA version and then upgrades to a newer one. Kept +# out of charts/openfga/ci/ because chart-testing only installs one version. replicaCount: 1 operator: diff --git a/.github/workflows/test.yml b/.github/workflows/test.yml index 7434585e..5efa259a 100644 --- a/.github/workflows/test.yml +++ b/.github/workflows/test.yml @@ -111,9 +111,7 @@ jobs: done test "$ver" = "${OLD_VER}" - # Operator must scale the openfga Deployment from 0 to 1 ready replica. - # condition=Available alone returns true at 0/0 before scale-up; - # readyReplicas=1 is the load-bearing signal. + # The pod stays NotReady until the migration has run on the new database. kubectl wait deployment/"$REL" -n "$NS" \ --for=jsonpath='{.status.readyReplicas}'=1 --timeout=3m diff --git a/charts/openfga-operator/templates/role.yaml b/charts/openfga-operator/templates/role.yaml index eae4d873..e9e851f4 100644 --- a/charts/openfga-operator/templates/role.yaml +++ b/charts/openfga-operator/templates/role.yaml @@ -8,7 +8,7 @@ metadata: rules: - apiGroups: ["apps"] resources: ["deployments"] - verbs: ["get", "list", "watch", "patch"] + verbs: ["get", "list", "watch"] - apiGroups: ["apps"] resources: ["deployments/status"] verbs: ["patch"] diff --git a/charts/openfga-operator/values.schema.json b/charts/openfga-operator/values.schema.json index 324465ff..07fbb49c 100644 --- a/charts/openfga-operator/values.schema.json +++ b/charts/openfga-operator/values.schema.json @@ -70,7 +70,7 @@ }, "activeDeadlineSeconds": { "type": "integer", - "minimum": 1 + "minimum": 0 }, "ttlSecondsAfterFinished": { "type": "integer", diff --git a/charts/openfga-operator/values.yaml b/charts/openfga-operator/values.yaml index 8e29275c..be5f05f6 100644 --- a/charts/openfga-operator/values.yaml +++ b/charts/openfga-operator/values.yaml @@ -59,7 +59,9 @@ migrationJob: # -- Number of pod failures before a migration Job is considered failed. backoffLimit: 3 # -- Maximum wall-clock seconds a migration Job can run before being terminated. - activeDeadlineSeconds: 300 + # 0 disables the deadline. A migration that is cut off, such as an index build + # on a large table, has to start over on the next attempt. + activeDeadlineSeconds: 0 # -- Seconds to keep completed/failed Job pods for log inspection before garbage collection. ttlSecondsAfterFinished: 300 diff --git a/charts/openfga/ci/operator-mode-values.yaml b/charts/openfga/ci/operator-mode-values.yaml index b85a6af2..02cb06b4 100644 --- a/charts/openfga/ci/operator-mode-values.yaml +++ b/charts/openfga/ci/operator-mode-values.yaml @@ -1,16 +1,8 @@ -# Exercises operator-managed mode end-to-end via chart-testing. -# -# The openfga-operator subchart auto-installs (conditional dependency on -# operator.enabled). With the memory datastore, the chart starts the -# Deployment at replicas=1 immediately, so `helm test` runs without racing -# the operator's reconcile loop. Migration is skipped (memory engine), but -# the rest of the wiring is exercised: subchart resolution, operator RBAC, -# pod/SA/annotation rendering, and the operator pod actually running and -# reconciling against the openfga Deployment in its release namespace. -# -# Postgres + operator (which exercises the migration Job path) is left to -# a follow-up E2E test — it requires waiting for the operator to scale the -# Deployment up before `helm test` runs the gRPC probe. +# Installs the openfga-operator subchart next to OpenFGA through chart-testing: +# subchart resolution, operator RBAC and the operator pod starting up. The +# memory datastore needs no migration, so the Deployment is not opted in to +# operator migrations; the Postgres migration path is covered by the operator +# E2E step in .github/workflows/test.yml. operator: enabled: true diff --git a/charts/openfga/templates/NOTES.txt b/charts/openfga/templates/NOTES.txt index 94bd3c71..ed84cd71 100644 --- a/charts/openfga/templates/NOTES.txt +++ b/charts/openfga/templates/NOTES.txt @@ -1,20 +1,17 @@ -{{- if and .Values.operator.enabled .Values.migration.enabled .Values.datastore.applyMigrations (has .Values.datastore.engine (list "postgres" "mysql")) }} -NOTE: operator-managed migration is enabled. The chart does not set spec.replicas -in this mode — the operator scales the Deployment to the desired count once the -migration Job succeeds. A fresh install begins at the Kubernetes default of one -replica (held NotReady by the readiness gate until migration finishes); an -upgrade preserves the live replica count while the migration runs. +{{- if include "openfga.operatorMigrations" . }} +NOTE: database migrations are run by the openfga-operator. Whenever the OpenFGA +image changes it runs the {{ include "openfga.fullname" . }}-migrate Job and records the migrated +version in the {{ include "openfga.fullname" . }}-migration-status ConfigMap. On a new database the +OpenFGA pods stay NotReady until the first migration completes. -If pods don't appear within ~2 minutes, check the operator and the migration -Job: +If the pods do not become ready, check the operator and the migration Job: - kubectl get deployment -A -l app.kubernetes.io/name=openfga-operator kubectl logs -n {{ .Release.Namespace }} -l app.kubernetes.io/name=openfga-operator --tail=100 kubectl get job/{{ include "openfga.fullname" . }}-migrate -n {{ .Release.Namespace }} -o yaml kubectl describe deployment/{{ include "openfga.fullname" . }} -n {{ .Release.Namespace }} -A `MigrationFailed` condition on the Deployment indicates the migration Job -failed; the operator will retry every 60s once the underlying issue is fixed. +A `MigrationFailed` condition on the Deployment means the migration Job failed; +the operator retries it every 60s. {{ end -}} 1. Get the application URL by running these commands: diff --git a/charts/openfga/templates/_helpers.tpl b/charts/openfga/templates/_helpers.tpl index cc50e03d..90d64bcd 100644 --- a/charts/openfga/templates/_helpers.tpl +++ b/charts/openfga/templates/_helpers.tpl @@ -85,6 +85,15 @@ Create the name of the migration service account to use (operator mode only) {{- end }} {{- end }} +{{/* +Return true if the openfga-operator runs the database migrations for this release +*/}} +{{- define "openfga.operatorMigrations" -}} +{{- if and .Values.operator.enabled .Values.migration.enabled .Values.datastore.applyMigrations (has .Values.datastore.engine (list "postgres" "mysql")) -}} +true +{{- end -}} +{{- end -}} + {{/* Return true if a secret object should be created */}} diff --git a/charts/openfga/templates/deployment.yaml b/charts/openfga/templates/deployment.yaml index 513eded8..80d07da5 100644 --- a/charts/openfga/templates/deployment.yaml +++ b/charts/openfga/templates/deployment.yaml @@ -4,13 +4,12 @@ metadata: name: {{ include "openfga.fullname" . }} labels: {{- include "openfga.labels" . | nindent 4 }} - {{- $hasOperatorAnnotations := and .Values.operator.enabled .Values.migration.enabled .Values.datastore.applyMigrations }} - {{- if or $hasOperatorAnnotations .Values.annotations }} + {{- $operatorMigrations := include "openfga.operatorMigrations" . }} + {{- if or $operatorMigrations .Values.annotations }} annotations: - {{- if $hasOperatorAnnotations }} + {{- if $operatorMigrations }} openfga.dev/migration-enabled: "true" openfga.dev/container-name: "{{ .Chart.Name }}" - openfga.dev/desired-replicas: '{{ ternary 1 .Values.replicaCount (eq .Values.datastore.engine "memory") }}' {{- if or .Values.migration.serviceAccount.create .Values.migration.serviceAccount.name }} openfga.dev/migration-service-account: '{{ include "openfga.migrationServiceAccountName" . }}' {{- end }} @@ -20,14 +19,7 @@ metadata: {{- end }} {{- end }} spec: - {{- if $hasOperatorAnnotations }} - {{- if .Values.autoscaling.enabled }} - {{- fail "operator.enabled and autoscaling.enabled cannot both be true" }} - {{- end }} - {{- /* Operator mode: omit spec.replicas. The operator owns it (scales up after - the migration Job completes) and GitOps controllers treat an absent field - as unmanaged, so neither fights the other. See ADR-002. */ -}} - {{- else if not .Values.autoscaling.enabled }} + {{- if not .Values.autoscaling.enabled }} replicas: {{ ternary 1 .Values.replicaCount (eq .Values.datastore.engine "memory") }} {{- end }} selector: diff --git a/charts/openfga/templates/serviceaccount.yaml b/charts/openfga/templates/serviceaccount.yaml index f732c46e..bc3e4aea 100644 --- a/charts/openfga/templates/serviceaccount.yaml +++ b/charts/openfga/templates/serviceaccount.yaml @@ -10,7 +10,7 @@ metadata: {{- toYaml . | nindent 4 }} {{- end }} {{- end }} -{{- if and .Values.operator.enabled .Values.migration.enabled .Values.migration.serviceAccount.create }} +{{- if and (include "openfga.operatorMigrations" .) .Values.migration.serviceAccount.create }} --- apiVersion: v1 kind: ServiceAccount diff --git a/charts/openfga/tests/operator_mode_serviceaccount_test.yaml b/charts/openfga/tests/operator_mode_serviceaccount_test.yaml index cbeab1a0..304f4ae7 100644 --- a/charts/openfga/tests/operator_mode_serviceaccount_test.yaml +++ b/charts/openfga/tests/operator_mode_serviceaccount_test.yaml @@ -6,6 +6,7 @@ tests: set: operator.enabled: true migration.enabled: true + datastore.engine: postgres migration.serviceAccount.create: true serviceAccount.create: true asserts: @@ -31,6 +32,7 @@ tests: set: operator.enabled: true migration.enabled: true + datastore.engine: postgres migration.serviceAccount.create: false migration.serviceAccount.name: external-sa serviceAccount.create: true @@ -42,6 +44,7 @@ tests: set: operator.enabled: true migration.enabled: true + datastore.engine: postgres migration.serviceAccount.create: true migration.serviceAccount.annotations: eks.amazonaws.com/role-arn: "arn:aws:iam::123456789012:role/openfga-migrator" @@ -56,6 +59,7 @@ tests: set: operator.enabled: true migration.enabled: true + datastore.engine: postgres migration.serviceAccount.create: true migration.serviceAccount.name: my-migrator serviceAccount.create: true @@ -64,3 +68,13 @@ tests: path: metadata.name value: my-migrator documentIndex: 1 + + - it: should not render migration service account for the memory datastore + set: + operator.enabled: true + migration.enabled: true + migration.serviceAccount.create: true + serviceAccount.create: true + asserts: + - hasDocuments: + count: 1 diff --git a/charts/openfga/tests/operator_mode_test.yaml b/charts/openfga/tests/operator_mode_test.yaml index bdffcf03..cefcbd08 100644 --- a/charts/openfga/tests/operator_mode_test.yaml +++ b/charts/openfga/tests/operator_mode_test.yaml @@ -7,15 +7,14 @@ tests: set: operator.enabled: true migration.enabled: true - replicaCount: 3 datastore.engine: postgres asserts: - equal: path: metadata.annotations["openfga.dev/migration-enabled"] value: "true" - equal: - path: metadata.annotations["openfga.dev/desired-replicas"] - value: "3" + path: metadata.annotations["openfga.dev/container-name"] + value: openfga - equal: path: metadata.annotations["openfga.dev/migration-service-account"] value: RELEASE-NAME-openfga-migration @@ -28,19 +27,18 @@ tests: asserts: - isNull: path: metadata.annotations["openfga.dev/migration-enabled"] - - isNull: - path: metadata.annotations["openfga.dev/desired-replicas"] + - equal: + path: metadata.annotations.custom + value: value - - it: should set desired-replicas to 1 for memory datastore + - it: should not set operator annotations for the memory datastore set: operator.enabled: true migration.enabled: true - replicaCount: 5 datastore.engine: memory asserts: - - equal: - path: metadata.annotations["openfga.dev/desired-replicas"] - value: "1" + - isNull: + path: metadata.annotations - it: should use custom migration service account name when set set: @@ -64,27 +62,30 @@ tests: path: metadata.annotations["openfga.dev/migration-service-account"] # --- Replica count --- - # Operator mode omits spec.replicas so the operator owns it and GitOps treats it - # as unmanaged. See ADR-002. - - it: should omit replicas in operator mode with a database datastore + # The operator never changes the replica count, so it renders as in legacy mode. + - it: should set replicas to replicaCount in operator mode set: operator.enabled: true migration.enabled: true replicaCount: 3 datastore.engine: postgres asserts: - - isNull: + - equal: path: spec.replicas + value: 3 - - it: should omit replicas in operator mode with a memory datastore + - it: should leave replicas to the autoscaler in operator mode set: operator.enabled: true migration.enabled: true - replicaCount: 5 - datastore.engine: memory + autoscaling.enabled: true + datastore.engine: postgres asserts: - isNull: path: spec.replicas + - equal: + path: metadata.annotations["openfga.dev/migration-enabled"] + value: "true" - it: should set replicas to replicaCount when operator is disabled set: @@ -111,17 +112,6 @@ tests: path: spec.replicas value: 4 - # --- Autoscaling conflict --- - - it: should fail when operator and autoscaling are both enabled - set: - operator.enabled: true - migration.enabled: true - autoscaling.enabled: true - datastore.engine: postgres - asserts: - - failedTemplate: - errorMessage: "operator.enabled and autoscaling.enabled cannot both be true" - # --- initContainers gating --- - it: should not render migration initContainers when operator is enabled set: diff --git a/charts/openfga/values.yaml b/charts/openfga/values.yaml index 9193c9d2..94be1ded 100644 --- a/charts/openfga/values.yaml +++ b/charts/openfga/values.yaml @@ -396,7 +396,7 @@ operator: openfga-operator: {} # migrationJob: # backoffLimit: 3 - # activeDeadlineSeconds: 300 + # activeDeadlineSeconds: 0 # ttlSecondsAfterFinished: 300 # leaderElection: # enabled: true @@ -416,7 +416,6 @@ migration: serviceAccount: # -- Create a dedicated service account for migration Jobs. # The migration Job inherits env vars (including secretKeyRef) from the OpenFGA container. - # If your datastore secret has RBAC restrictions, ensure this service account can read it. create: true # -- Annotations to add to the migration service account. # Use this to attach cloud IAM roles (e.g., eks.amazonaws.com/role-arn) for DDL permissions. @@ -424,6 +423,7 @@ migration: # -- The name of the migration service account. # If not set and create is true, defaults to {fullname}-migration. name: "" + ## Example: Deploy a PostgreSQL instance for dev/test using official Docker images. ## For production, use a managed database service or an operator like CloudnativePG. ## Configure the chart to use the secret: diff --git a/docs/adr/001-adopt-openfga-operator.md b/docs/adr/001-adopt-openfga-operator.md index cf8c44c5..63ee980f 100644 --- a/docs/adr/001-adopt-openfga-operator.md +++ b/docs/adr/001-adopt-openfga-operator.md @@ -53,7 +53,7 @@ The OpenFGA Helm chart currently handles all lifecycle concerns — deployment, We will build an **OpenFGA Kubernetes Operator** that handles: -1. **Database migration orchestration** (Stage 1) — replacing Helm hooks, the `k8s-wait-for` init container, and shared ServiceAccount with operator-managed migration Jobs and deployment readiness gating. +1. **Database migration orchestration** (Stage 1) — replacing Helm hooks, the `k8s-wait-for` init container, and shared ServiceAccount with operator-managed migration Jobs. 2. **Declarative store lifecycle management** (Stages 2-4) — exposing `FGAStore`, `FGAModel`, and `FGATuples` CRDs for GitOps-native authorization configuration. @@ -80,7 +80,7 @@ Stage 1 has shipped on the `feat/operator-migration` branch. Stages 2-4 are plan - Operator Go project under `/operator/`, built with `controller-runtime` and kubebuilder scaffolding - Operator packaged as a Helm subchart (`charts/openfga-operator/`) and wired into the main chart via a `condition: operator.enabled` dependency - `operator.enabled` values toggle (default `false`) that gates all operator-managed behavior -- Migration reconciler (`migration_controller.go`) that orchestrates migration Jobs and gates Deployment readiness when the operator is enabled +- Migration reconciler (`migration_controller.go`) that runs migration Jobs when the operator is enabled - Separate migration ServiceAccount with IAM-annotation support (`openfga.migrationServiceAccountName` helper), created when the operator is enabled ### Deferred to later stages diff --git a/docs/adr/002-operator-managed-migrations.md b/docs/adr/002-operator-managed-migrations.md index 0f938f65..5f35a0cb 100644 --- a/docs/adr/002-operator-managed-migrations.md +++ b/docs/adr/002-operator-managed-migrations.md @@ -88,49 +88,43 @@ The operator runs a **migration controller** that reconciles the OpenFGA Deploym │ └── ttlSecondsAfterFinished: 300 │ │ 5. Watch Job until succeeded │ │ 6. Update ConfigMap → "version: v1.14.0" │ -│ 7. Scale Deployment to desired replicas │ -│ (fresh install: default 1 → N; upgrade: unchanged) │ -│ 8. New pods pass readiness, serve requests │ └──────────────────────────────────────────────────────────┘ ``` **Key design decisions within this approach:** -#### Zero-downtime upgrades via omitted replicas and readiness gating +#### The operator only runs migrations -In operator mode the chart **omits `spec.replicas` entirely** rather than rendering a fixed number. The operator owns the replica count: it scales the Deployment to `openfga.dev/desired-replicas` once the migration Job succeeds. Omitting the field (the same treatment an HPA-managed Deployment gets) means a GitOps controller sees no declared replica count and leaves it unmanaged, so it never fights the operator over the value — no `ignoreDifferences` or field-ownership patch is required. +The operator creates Jobs and records their outcome; it never changes the Deployment's replica count or pod template. The chart renders `spec.replicas` exactly as in legacy mode (or leaves it to an HPA), so `kubectl scale`, autoscalers and GitOps tools behave the same whether or not the operator is enabled. -On **fresh install**, no `spec.replicas` is set, so Kubernetes applies its default of one replica. That pod starts before the migration has run and is held `NotReady` by OpenFGA's readiness gate (see below), so it serves no traffic. The operator runs the migration Job and then scales the Deployment to the desired replica count. +Readiness comes from OpenFGA itself: `IsReady()` reports `NOT_SERVING` while the schema revision is below `MinimumSupportedDatastoreSchemaRevision` (4 since v1.3.x). On a **fresh install** the database is empty, so every pod stays `NotReady` until the first migration Job completes, and `helm install --wait` returns once it has. On an **upgrade** the existing schema already meets that minimum, so new pods pass readiness right away and serve on the previous schema while the Job applies the newer migrations. This relies on OpenFGA migrations being backward compatible, which is also what the Helm hook flow has always done: its init container sees the previous release's completed hook Job and lets new pods start before the new hook runs. -On **upgrade**, the field is still absent from the rendered manifest, so Kubernetes preserves the live replica count and starts a rolling update with the new image. OpenFGA has a **built-in schema version gate**: on startup, each instance calls `IsReady()` which checks the database schema revision against `MinimumSupportedDatastoreSchemaRevision` (via goose). If the schema is behind, the gRPC health endpoint returns `NOT_SERVING`, the readiness probe fails, and Kubernetes does not route traffic to the pod. Old pods continue serving on the migrated schema (OpenFGA migrations are additive/backward-compatible — this is how the existing Helm hook flow has operated for years with rolling updates). Once the operator's migration Job completes, new pods pass readiness and the rolling update proceeds. - -This matches the existing zero-downtime behavior of the non-operator chart. - -**Rejected alternative — pin `replicas: 0` and read the live count via `lookup`:** the chart could render `replicas: 0` on fresh install and use Helm's `lookup` to preserve the live count on upgrade. This is rejected for two reasons. A GitOps controller that reconciles the manifest drives replicas back to 0 on every sync and fights the operator, causing a full outage. And `lookup` returns empty under `helm template` and `--dry-run=client`, so the manifest silently renders `replicas: 0` in CI and dry runs. Omitting the field avoids both problems and needs no cluster lookup. +**Rejected alternative — let the operator own the replica count:** the chart could omit `spec.replicas` (or render 0) and have the operator scale the Deployment up once the migration succeeds. Testing this showed three problems: switching an existing release to operator mode removes the field, so both Helm's three-way merge and server-side apply reset the Deployment to one replica until the migration finishes; `kubectl scale` and HPAs are overridden by the operator; and the scale-up buys nothing on upgrades, where the readiness check does not hold pods back. #### Version tracking via ConfigMap A ConfigMap (`openfga-migration-status`) records the last successfully migrated version. The operator compares this to the Deployment's image tag to determine if migration is needed. This is: - Simple to inspect (`kubectl get configmap openfga-migration-status -o yaml`) - Survives operator restarts -- Can be manually deleted to force re-migration +- Can be manually deleted to force re-migration (once the previous migration Job has been cleaned up) #### Separate ServiceAccount for migrations -The operator creates a dedicated `openfga-migrator` ServiceAccount for migration Jobs. Users can annotate it with cloud IAM roles that grant DDL permissions, while the runtime ServiceAccount retains only CRUD permissions. +The chart creates a dedicated `{fullname}-migration` ServiceAccount that the operator uses for migration Jobs. Users can annotate it with cloud IAM roles that grant DDL permissions, while the runtime ServiceAccount retains only CRUD permissions. #### Migration Job is a regular resource -The Job created by the operator has no Helm hook annotations. It is a standard Kubernetes Job, visible to ArgoCD, FluxCD, and all Kubernetes tooling. It has an owner reference to the operator's managed resource for proper garbage collection. +The Job created by the operator has no Helm hook annotations. It is a standard Kubernetes Job, visible to ArgoCD, FluxCD, and all Kubernetes tooling. It has an owner reference to the OpenFGA Deployment, so it is garbage collected with it. #### Failure handling | Failure | Behavior | |---------|----------| -| Job fails | Operator sets `MigrationFailed` on the Deployment and does not scale it. New pods stay `NotReady` behind the readiness gate; existing pods keep serving. | -| Job hangs | `activeDeadlineSeconds` (default 300s) kills it. Operator sees failure. | -| Operator crashes | On restart, re-reads ConfigMap and Job status. Resumes from where it left off. | -| Database unreachable | Job fails to connect. After exhausting `backoffLimit`, operator deletes the failed Job, sets a `retry-after` annotation, and recreates a fresh Job after a fixed 60-second cooldown. Cycle repeats until the database becomes available. | +| Job fails | Operator sets `MigrationFailed` on the Deployment, keeps the failed Job for 60 seconds so its logs can be read, then replaces it. On a fresh database the pods stay `NotReady`; on an upgrade they keep serving on the previous schema. | +| Job pod never starts | A bad secret reference, image pull error or unschedulable pod never fails the Job. Once the Deployment's pod template changes (the fix rolls out), the operator rebuilds a Job whose pod is not running. | +| Job hangs | No deadline by default, like the Helm hook Job. `activeDeadlineSeconds` can be set, but a migration cut off halfway (an index build, a MySQL table rebuild) starts over on the next attempt. | +| Operator crashes | On restart, re-reads the ConfigMap and Job status and resumes. The retry delay is measured from the failed Job's condition, so it survives restarts. | +| Database unreachable | Job fails to connect. After exhausting `backoffLimit` the cycle above repeats until the database becomes available. | ### Sequence Comparison @@ -157,12 +151,11 @@ Problems: ArgoCD skips step 4. FluxCD deletes Job in step 4. `--wait` deadlocks helm install ├── Create ServiceAccount (runtime), ServiceAccount (migrator) ├── Create Secret, Service - ├── Create Deployment (no spec.replicas, no init containers) + ├── Create Deployment (no init containers) ├── Create Operator Deployment └── [Helm is done — all resources are regular, no hooks] -(Kubernetes starts the Deployment at its default of 1 replica; that pod -is held NotReady by the readiness gate until the migration completes.) +(The OpenFGA pods start but stay NotReady: the database has no schema yet.) Operator starts: ├── Detects Deployment image version @@ -171,30 +164,25 @@ Operator starts: │ └── Uses openfga-migrator ServiceAccount │ └── Runs openfga migrate → succeeds ├── Creates ConfigMap with migrated version - └── Scales Deployment to 3 replicas → pods pass readiness + └── Pods pass readiness ``` **After (operator-managed, upgrade with new image):** ```text helm upgrade - ├── no spec.replicas in manifest → Kubernetes keeps the live count (3) ├── Patches Deployment with new image tag ├── Kubernetes starts rolling update - │ ├── New pods (v1.14) start → schema is behind → - │ │ readiness fails (gRPC NOT_SERVING) → no traffic routed - │ └── Old pods (v1.13) continue serving traffic + │ └── New pods (v1.14) pass readiness on the previous schema └── [Helm is done] Operator reconciles: ├── Detects image version differs from ConfigMap ├── Creates Job/openfga-migrate → runs migration - ├── Updates ConfigMap → "version: v1.14.0" - └── New pods pass readiness → rolling update completes - (operator does NOT scale to zero — zero downtime) + └── Updates ConfigMap → "version: v1.14.0" ``` -No hooks. No init containers. No `k8s-wait-for`. No downtime on upgrade. All resources are regular Kubernetes objects. +No hooks. No init containers. No `k8s-wait-for`. All resources are regular Kubernetes objects. ### What Changes in the Helm Chart @@ -209,7 +197,7 @@ Nothing is deleted outright — every change is gated on `operator.enabled` so t | `values.yaml`: `initContainer.*` | Unused — `k8s-wait-for` not deployed | | `values.yaml`: `datastore.migrationType`, `datastore.waitForMigrations` | Unused — operator always uses a Job and handles ordering | | `values.yaml`: `migrate.annotations` | Unused — no Helm hooks | -| Deployment migration init containers | Skipped — operator manages readiness via replica scaling | +| Deployment migration init containers | Skipped — OpenFGA's readiness check holds pods until the schema is migrated | **Added (active only when `operator.enabled: true`):** @@ -237,10 +225,10 @@ Users on `operator.enabled: false` (the default) see identical rendered output t ### Negative - **Operator is a new runtime dependency** — if the operator pod is unavailable, migrations don't run (but existing running pods are unaffected) -- **Replica count is unmanaged by the chart in operator mode** — because `spec.replicas` is omitted, the rendered manifest no longer declares a desired count; the operator (and, on first install, the Kubernetes default of 1) determines it. A reader inspecting only the chart output cannot see the running replica count. - **Two upgrade paths to document** — `operator.enabled: true` (new) vs `operator.enabled: false` (legacy) ### Risks -- **Readiness gate relies on OpenFGA's built-in schema check** — the zero-downtime upgrade model depends on `MinimumSupportedDatastoreSchemaRevision` in `pkg/storage/sqlcommon/sqlcommon.go` causing `NOT_SERVING` when the schema is behind. If a future OpenFGA release removes or weakens this check, new pods could serve traffic against an unmigrated schema. This coupling should be documented and monitored across OpenFGA releases. -- **ConfigMap as state store** — if the ConfigMap is accidentally deleted, the operator re-runs migration (which is safe — `openfga migrate` is idempotent). This is a feature, not a bug, but should be documented. +- **Readiness relies on OpenFGA's schema check** — pods on a fresh database are held back only by `MinimumSupportedDatastoreSchemaRevision` in `pkg/storage/sqlcommon/sqlcommon.go`, and upgrades rely on each release working against the previous schema. Both are OpenFGA guarantees the Helm hook flow already depended on. +- **Migrations run as soon as the image changes** — as with the hook Job, nothing drains traffic first. Some migrations, such as MySQL's `008_collate_identifiers` in v1.18.0, block writes while tables are rebuilt; OpenFGA's runbook recommends draining traffic for those, which stays a manual step. +- **ConfigMap as state store** — if the ConfigMap is accidentally deleted, the operator records the version again from the completed Job while it exists, or re-runs the migration once it has been cleaned up (which is safe — `openfga migrate` is idempotent). diff --git a/operator/README.md b/operator/README.md index 480b740b..6a306d93 100644 --- a/operator/README.md +++ b/operator/README.md @@ -11,10 +11,9 @@ This is **Stage 1** of the operator — focused solely on migration orchestratio - Creates a migration Job running `openfga migrate` - Waits for the Job to complete - Updates the ConfigMap with the new version - - Scales the Deployment to the desired replica count (`openfga.dev/desired-replicas`) -3. On failure, a `MigrationFailed` condition is set on the Deployment and the desired replica count is not applied +3. On failure, a `MigrationFailed` condition is set on the Deployment. The failed Job is kept for 60 seconds so its logs can be inspected, then replaced with a new one. -The operator never scales the Deployment to 0. A pod that starts before the migration completes is held `NotReady` by OpenFGA's readiness gate on `MinimumSupportedDatastoreSchemaRevision`, so it won't serve traffic against an unmigrated schema. +The operator never changes the Deployment's replica count or pod template. On a new database, OpenFGA's readiness check (`MinimumSupportedDatastoreSchemaRevision`) keeps pods `NotReady` until the first migration has run. On an upgrade the existing schema already meets that minimum, so new pods serve on it while the Job applies the newer migrations, which is the same behaviour as the Helm hook flow. ## Prerequisites @@ -56,9 +55,9 @@ Integration test values and instructions are in [`tests/`](tests/). Three scenar | Scenario | Values File | What It Tests | |----------|-------------|---------------| -| Happy path | `tests/values-happy-path.yaml` | Full lifecycle: Postgres up, migration succeeds, OpenFGA scales to 3/3 | +| Happy path | `tests/values-happy-path.yaml` | Full lifecycle: Postgres up, migration succeeds, OpenFGA ready at 3/3 | | DB outage & recovery | `tests/values-db-outage.yaml` | Postgres starts at 0 replicas; scale it up later to verify self-healing | -| No database | `tests/values-no-db.yaml` | Permanent failure: operator retries without crashing; the app pod stays NotReady (0/1) | +| No database | `tests/values-no-db.yaml` | Permanent failure: operator retries without crashing; the app pods stay NotReady (0/3) | Quick start: @@ -96,7 +95,7 @@ operator/ │ └── controller/ │ ├── migration_controller.go # Reconciliation loop │ ├── migration_controller_test.go # Unit tests -│ └── helpers.go # Job builder, scaling, ConfigMap helpers +│ └── helpers.go # Job builder, status ConfigMap helpers ├── Dockerfile # Multi-stage build (distroless runtime) ├── Makefile ├── go.mod @@ -113,8 +112,8 @@ The operator accepts the following flags: | `--watch-namespace` | `""` | Namespace to watch for OpenFGA Deployments. Defaults to the operator pod's own namespace (via `POD_NAMESPACE` env var). The chart binds namespaced RBAC in the configured watch namespace, so the operator may run in a different namespace when needed. | | `--metrics-bind-address` | `:8080` | Address the Prometheus metrics endpoint binds to. Change only if the default port conflicts with other containers in the pod. | | `--health-probe-bind-address` | `:8081` | Address the Kubernetes liveness and readiness probe endpoints bind to. Change only if the default port conflicts. | -| `--backoff-limit` | `3` | Number of times a migration Job's pod can fail before the Job is considered failed. After hitting this limit the operator deletes the Job, sets a `MigrationFailed` condition on the Deployment, and retries after a 60-second cooldown. | -| `--active-deadline-seconds` | `300` | Maximum wall-clock seconds a migration Job can run before Kubernetes terminates it. Prevents stuck migrations from blocking the pipeline indefinitely. Increase for very large databases. | +| `--backoff-limit` | `3` | Number of times a migration Job's pod can fail before the Job is considered failed. The operator then sets a `MigrationFailed` condition on the Deployment and replaces the Job 60 seconds after it failed. | +| `--active-deadline-seconds` | `0` | Maximum wall-clock seconds a migration Job can run before Kubernetes terminates it. `0` means no deadline. A deadline cuts off long migrations, such as index builds or MySQL table rebuilds on large tables, which then start over on the next attempt. | | `--ttl-seconds-after-finished` | `300` | Seconds Kubernetes keeps a completed or failed Job (and its pods) before garbage-collecting them, giving you time to inspect logs. | When deployed via the Helm subchart, these are configured through `values.yaml`. See `charts/openfga-operator/values.yaml` for all available options. @@ -125,14 +124,13 @@ The operator reads these annotations from the OpenFGA Deployment: | Annotation | Description | |------------|-------------| -| `openfga.dev/migration-enabled` | Must be `"true"` for the operator to manage migrations. Deployments without this annotation are ignored. Set by the Helm chart when `operator.enabled`, `migration.enabled`, and `datastore.applyMigrations` are all true. | -| `openfga.dev/desired-replicas` | The replica count the operator scales the Deployment to once migration succeeds. Set by the Helm chart. | +| `openfga.dev/migration-enabled` | Must be `"true"` for the operator to manage migrations. Deployments without this annotation are ignored. Set by the Helm chart when `operator.enabled`, `migration.enabled`, and `datastore.applyMigrations` are true and the datastore is Postgres or MySQL. | +| `openfga.dev/container-name` | The OpenFGA container in the pod spec. Defaults to `openfga`. | | `openfga.dev/migration-service-account` | The ServiceAccount to use for migration Jobs. Defaults to the Deployment's SA. | ## Limitations -- **Migrations key only on the image tag:** The operator compares the container image tag (or digest) to the `{name}-migration-status` ConfigMap. A mutable tag like `latest`, or a tag reused for a new build, is not seen as a change, so the migration is skipped — use immutable tags (e.g. `v1.14.0`) or pin by digest. A migration-needing change that keeps the same image — for example repointing `datastore.uri` at a different or restored database — also won't trigger a Job; the readiness gate holds the new pod `NotReady`, but you must migrate manually (bump the image or delete the status ConfigMap). -- **Migration-specific volumes:** The legacy Helm chart values `migrate.extraVolumes` and `migrate.extraVolumeMounts` have no effect in operator mode. The operator inherits volumes and mounts from the main Deployment pod spec. If you need additional volumes for migrations (e.g., CA bundles or TLS certs), add them to the top-level `extraVolumes` and `extraVolumeMounts` values instead. -- **Single-container migration Job:** The Job runs one container (`openfga migrate`) with the main container's env, volumes, and scheduling. It injects no sidecars or extra init containers, so databases reached through a sidecar proxy (Cloud SQL Auth Proxy, AlloyDB) aren't supported for operator-managed migrations — a proxy that doesn't exit on its own (e.g. an Istio sidecar) would keep the Job pod running and stop the Job from completing. Connect to such databases directly instead. -- **`envFrom` datastore detection:** The memory-datastore check inspects only the explicit `env` entries on the container. If `OPENFGA_DATASTORE_ENGINE` is supplied via `envFrom` (a ConfigMap or Secret), the operator cannot read the value and will attempt a migration Job that a memory datastore does not need. The Helm chart sets this variable inline, so chart-managed installs are unaffected. -- **GitOps and `spec.replicas`:** The operator owns the replica count — it scales to `openfga.dev/desired-replicas` after the migration Job completes (see [ADR-002](../docs/adr/002-operator-managed-migrations.md)). To avoid fighting a GitOps controller over that field, the chart omits `spec.replicas` in operator mode, the same as an HPA-managed Deployment. Argo CD and Flux treat an absent field as unmanaged, so no `ignoreDifferences` or field-ownership patch is needed. A fresh install starts at the Kubernetes default of one replica, held `NotReady` by the readiness gate until the migration finishes, after which the operator scales to the desired count. +- **Migrations key only on the image tag:** The operator compares the container image tag (or digest) to the `{name}-migration-status` ConfigMap. A mutable tag like `latest`, or a tag reused for a new build, is not seen as a change, so the migration is skipped — use immutable tags (e.g. `v1.14.0`) or pin by digest. A migration-needing change that keeps the same image — for example repointing `datastore.uri` at a different or restored database — also won't trigger a Job; delete the status ConfigMap (and the `{name}-migrate` Job, if it still exists) to run the migration again. +- **Legacy migration values:** `migrate.*` (extra volumes and mounts, init containers, sidecars, annotations, labels, timeout) and `datastore.migrations.resources` only apply to the legacy Helm hook Job. The operator's Job copies the OpenFGA container's image, env, volumes, resources, security context and scheduling instead, so put anything the migration needs (e.g. CA bundles) in the top-level `extraVolumes`, `extraVolumeMounts` and `extraEnvVars`. +- **Single-container migration Job:** The Job runs one container (`openfga migrate`) and injects no sidecars or extra init containers, so databases reached through a sidecar proxy (Cloud SQL Auth Proxy, AlloyDB) aren't supported for operator-managed migrations. A sidecar injected into every pod in the namespace that doesn't exit on its own (e.g. an Istio sidecar) keeps the Job pod running and stops the Job from completing. +- **One namespace per operator:** The operator reconciles every opted-in OpenFGA Deployment in its watch namespace. Operators installed by several releases in one namespace share a leader election lease, so only one of them is active at a time. diff --git a/operator/cmd/main.go b/operator/cmd/main.go index daf01e8e..94c65ce9 100644 --- a/operator/cmd/main.go +++ b/operator/cmd/main.go @@ -39,7 +39,7 @@ func main() { flag.StringVar(&metricsAddr, "metrics-bind-address", ":8080", "The address the metric endpoint binds to.") flag.StringVar(&healthProbeAddr, "health-probe-bind-address", ":8081", "The address the health probe endpoint binds to.") flag.IntVar(&backoffLimit, "backoff-limit", int(controller.DefaultBackoffLimit), "BackoffLimit for migration Jobs.") - flag.IntVar(&activeDeadline, "active-deadline-seconds", int(controller.DefaultActiveDeadlineSeconds), "ActiveDeadlineSeconds for migration Jobs.") + flag.IntVar(&activeDeadline, "active-deadline-seconds", int(controller.DefaultActiveDeadlineSeconds), "ActiveDeadlineSeconds for migration Jobs; 0 means no deadline.") flag.IntVar(&ttlAfterFinished, "ttl-seconds-after-finished", int(controller.DefaultTTLSecondsAfterFinished), "TTLSecondsAfterFinished for migration Jobs.") opts := zap.Options{Development: false} diff --git a/operator/internal/controller/helpers.go b/operator/internal/controller/helpers.go index 1c4f240c..a82f1799 100644 --- a/operator/internal/controller/helpers.go +++ b/operator/internal/controller/helpers.go @@ -2,9 +2,9 @@ package controller import ( "context" + "crypto/sha256" + "encoding/json" "fmt" - "regexp" - "strconv" "strings" "time" @@ -14,6 +14,7 @@ import ( metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" "k8s.io/utils/ptr" "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/controller/controllerutil" ) const ( @@ -26,63 +27,39 @@ const ( // Labels set on operator-managed resources (migration Jobs, status ConfigMaps). LabelManagedBy = "app.kubernetes.io/managed-by" - LabelVersion = "app.kubernetes.io/version" LabelManagedByValue = "openfga-operator" - // Annotations set on the Deployment by the Helm chart / operator. + // Annotations read from the Deployment (set by the Helm chart). AnnotationMigrationEnabled = "openfga.dev/migration-enabled" AnnotationContainerName = "openfga.dev/container-name" - AnnotationDesiredReplicas = "openfga.dev/desired-replicas" - AnnotationDesiredVersion = "openfga.dev/desired-version" AnnotationMigrationServiceAccount = "openfga.dev/migration-service-account" - AnnotationRetryAfter = "openfga.dev/migration-retry-after" - // Defaults for migration Job configuration. + // Annotations set on migration Jobs: the version the Job migrates to, and a + // hash of the pod template it was built from. + AnnotationDesiredVersion = "openfga.dev/desired-version" + AnnotationPodTemplateHash = "openfga.dev/pod-template-hash" + + // Defaults for migration Job configuration. An ActiveDeadlineSeconds of 0 + // leaves the Job without a deadline. DefaultBackoffLimit int32 = 3 - DefaultActiveDeadlineSeconds int64 = 300 + DefaultActiveDeadlineSeconds int64 = 0 DefaultTTLSecondsAfterFinished int32 = 300 ) -// immutableTag matches a fully-qualified semantic version tag (with an optional -// leading "v" and optional pre-release/build suffix), which is treated as -// immutable by convention. -var immutableTag = regexp.MustCompile(`^v?\d+\.\d+\.\d+([-+][0-9A-Za-z.-]+)?$`) - -// isMutableImageReference reports whether an image reference is not pinned to an -// immutable identifier. Digest references (@sha256:...) are immutable, and a -// full semantic version tag is treated as immutable by convention. Everything -// else — "latest", a floating "v1.14", a bare name — is mutable: the same -// reference can resolve to different images over time, so the operator cannot -// tell that a rebuilt image needs a migration. -func isMutableImageReference(image string) bool { - if strings.Contains(image, "@") { - return false - } - return !immutableTag.MatchString(extractImageTag(image)) -} - // extractImageTag returns the tag portion of a container image reference. // For "openfga/openfga:v1.14.0" it returns "v1.14.0". // For "openfga/openfga@sha256:abc..." it returns the digest. // If there is no tag or digest, it returns "latest". func extractImageTag(image string) string { - // Handle digest references. if idx := strings.LastIndex(image, "@"); idx != -1 { return image[idx+1:] } - // Handle tag references — be careful not to split on the port in a registry URL. - // Find the last '/' to isolate the image name from the registry. - lastSlash := strings.LastIndex(image, "/") - nameAndTag := image - if lastSlash != -1 { - nameAndTag = image[lastSlash+1:] - } - + // Only look for ":" after the last "/" so a registry port is not mistaken for a tag. + nameAndTag := image[strings.LastIndex(image, "/")+1:] if idx := strings.LastIndex(nameAndTag, ":"); idx != -1 { return nameAndTag[idx+1:] } - return "latest" } @@ -96,20 +73,14 @@ func migrationJobName(deploymentName string) string { return deploymentName + "-migrate" } -// findOpenFGAContainer finds the OpenFGA container in the Deployment's pod spec. -// It checks the openfga.dev/container-name annotation first, then looks for a -// container named "openfga". Returns an error if no containers exist or the -// target container is not found. +// findOpenFGAContainer returns the container named by the openfga.dev/container-name +// annotation, or the container named "openfga" when the annotation is absent. func findOpenFGAContainer(deployment *appsv1.Deployment) (*corev1.Container, error) { - containers := deployment.Spec.Template.Spec.Containers - if len(containers) == 0 { - return nil, fmt.Errorf("deployment %s/%s has no containers", deployment.Namespace, deployment.Name) - } - targetName := deployment.Annotations[AnnotationContainerName] if targetName == "" { targetName = "openfga" } + containers := deployment.Spec.Template.Spec.Containers for i := range containers { if containers[i].Name == targetName { return &containers[i], nil @@ -118,29 +89,30 @@ func findOpenFGAContainer(deployment *appsv1.Deployment) (*corev1.Container, err return nil, fmt.Errorf("container %q not found in deployment %s/%s", targetName, deployment.Namespace, deployment.Name) } -// buildMigrationJob constructs a migration Job for the given Deployment. -func buildMigrationJob( - deployment *appsv1.Deployment, - mainContainer *corev1.Container, - desiredVersion string, - backoffLimit int32, - activeDeadlineSeconds int64, - ttlSecondsAfterFinished int32, -) *batchv1.Job { - // Determine the migration service account. - migrationSA := deployment.Annotations[AnnotationMigrationServiceAccount] - if migrationSA == "" { - migrationSA = deployment.Spec.Template.Spec.ServiceAccountName +// ownerReference makes the Deployment the controller of a migration Job or +// status ConfigMap so both are garbage collected with it. BlockOwnerDeletion +// is left unset: it needs update on deployments/finalizers, which the +// operator is not granted. +func ownerReference(deployment *appsv1.Deployment) metav1.OwnerReference { + return metav1.OwnerReference{ + APIVersion: "apps/v1", + Kind: "Deployment", + Name: deployment.Name, + UID: deployment.UID, + Controller: ptr.To(true), } +} - // Sanitize version for use as a label value (must match [a-zA-Z0-9._-], max 63 chars). - // The full version is stored in an annotation for accurate comparison. - labelVersion := strings.ReplaceAll(desiredVersion, ":", "_") - if len(labelVersion) > 63 { - labelVersion = labelVersion[:63] +// buildMigrationJob constructs a Job that runs "openfga migrate" with the +// OpenFGA container's image, environment, volumes and scheduling. +func (r *MigrationReconciler) buildMigrationJob(deployment *appsv1.Deployment, container *corev1.Container, version string) *batchv1.Job { + podSpec := deployment.Spec.Template.Spec + serviceAccount := deployment.Annotations[AnnotationMigrationServiceAccount] + if serviceAccount == "" { + serviceAccount = podSpec.ServiceAccountName } - return &batchv1.Job{ + job := &batchv1.Job{ ObjectMeta: metav1.ObjectMeta{ Name: migrationJobName(deployment.Name), Namespace: deployment.Namespace, @@ -148,25 +120,13 @@ func buildMigrationJob( LabelPartOf: LabelPartOfValue, LabelComponent: "migration", LabelManagedBy: LabelManagedByValue, - LabelVersion: labelVersion, - }, - Annotations: map[string]string{ - AnnotationDesiredVersion: desiredVersion, - }, - OwnerReferences: []metav1.OwnerReference{ - { - APIVersion: "apps/v1", - Kind: "Deployment", - Name: deployment.Name, - UID: deployment.UID, - Controller: ptr.To(true), - }, }, + Annotations: map[string]string{AnnotationDesiredVersion: version}, + OwnerReferences: []metav1.OwnerReference{ownerReference(deployment)}, }, Spec: batchv1.JobSpec{ - BackoffLimit: ptr.To(backoffLimit), - ActiveDeadlineSeconds: ptr.To(activeDeadlineSeconds), - TTLSecondsAfterFinished: ptr.To(ttlSecondsAfterFinished), + BackoffLimit: ptr.To(r.BackoffLimit), + TTLSecondsAfterFinished: ptr.To(r.TTLSecondsAfterFinished), Template: corev1.PodTemplateSpec{ ObjectMeta: metav1.ObjectMeta{ Labels: map[string]string{ @@ -175,123 +135,67 @@ func buildMigrationJob( }, }, Spec: corev1.PodSpec{ - ServiceAccountName: migrationSA, + ServiceAccountName: serviceAccount, RestartPolicy: corev1.RestartPolicyNever, - ImagePullSecrets: deployment.Spec.Template.Spec.ImagePullSecrets, - SecurityContext: deployment.Spec.Template.Spec.SecurityContext, - Containers: []corev1.Container{ - { - Name: "migrate-database", - Image: mainContainer.Image, - ImagePullPolicy: mainContainer.ImagePullPolicy, - Args: []string{"migrate"}, - Env: mainContainer.Env, - EnvFrom: mainContainer.EnvFrom, - Resources: mainContainer.Resources, - VolumeMounts: mainContainer.VolumeMounts, - SecurityContext: mainContainer.SecurityContext, - }, - }, - // Inherit volumes and scheduling constraints from the parent Deployment. - Volumes: deployment.Spec.Template.Spec.Volumes, - NodeSelector: deployment.Spec.Template.Spec.NodeSelector, - Tolerations: deployment.Spec.Template.Spec.Tolerations, - Affinity: deployment.Spec.Template.Spec.Affinity, + ImagePullSecrets: podSpec.ImagePullSecrets, + SecurityContext: podSpec.SecurityContext, + Containers: []corev1.Container{{ + Name: "migrate-database", + Image: container.Image, + ImagePullPolicy: container.ImagePullPolicy, + Args: []string{"migrate"}, + Env: container.Env, + EnvFrom: container.EnvFrom, + Resources: container.Resources, + VolumeMounts: container.VolumeMounts, + SecurityContext: container.SecurityContext, + }}, + Volumes: podSpec.Volumes, + NodeSelector: podSpec.NodeSelector, + Tolerations: podSpec.Tolerations, + Affinity: podSpec.Affinity, }, }, }, } + if r.ActiveDeadlineSeconds > 0 { + job.Spec.ActiveDeadlineSeconds = ptr.To(r.ActiveDeadlineSeconds) + } + job.Annotations[AnnotationPodTemplateHash] = podTemplateHash(&job.Spec.Template) + return job } -// updateMigrationStatus creates or updates the migration-status ConfigMap. -func updateMigrationStatus( - ctx context.Context, - c client.Client, - deployment *appsv1.Deployment, - version string, - jobName string, -) error { - cmName := migrationConfigMapName(deployment.Name) - cm := &corev1.ConfigMap{ - ObjectMeta: metav1.ObjectMeta{ - Name: cmName, - Namespace: deployment.Namespace, - Labels: map[string]string{ - LabelPartOf: LabelPartOfValue, - LabelComponent: "migration", - LabelManagedBy: LabelManagedByValue, - }, - OwnerReferences: []metav1.OwnerReference{ - { - APIVersion: "apps/v1", - Kind: "Deployment", - Name: deployment.Name, - UID: deployment.UID, - Controller: ptr.To(true), - }, - }, - }, - Data: map[string]string{ - "version": version, - "migratedAt": time.Now().UTC().Format(time.RFC3339), - "jobName": jobName, - }, +func podTemplateHash(template *corev1.PodTemplateSpec) string { + b, err := json.Marshal(template) + if err != nil { + panic(err) // a PodTemplateSpec always marshals } + return fmt.Sprintf("%x", sha256.Sum256(b))[:16] +} - // Try to get existing ConfigMap first. - existing := &corev1.ConfigMap{} - err := c.Get(ctx, client.ObjectKeyFromObject(cm), existing) - if err != nil { - if client.IgnoreNotFound(err) != nil { - return fmt.Errorf("getting migration status ConfigMap: %w", err) +// updateMigrationStatus records the migrated version in the status ConfigMap. +func updateMigrationStatus(ctx context.Context, c client.Client, deployment *appsv1.Deployment, version, jobName string) error { + cm := &corev1.ConfigMap{ObjectMeta: metav1.ObjectMeta{ + Name: migrationConfigMapName(deployment.Name), + Namespace: deployment.Namespace, + }} + _, err := controllerutil.CreateOrUpdate(ctx, c, cm, func() error { + cm.Labels = map[string]string{ + LabelPartOf: LabelPartOfValue, + LabelComponent: "migration", + LabelManagedBy: LabelManagedByValue, } - // ConfigMap doesn't exist — create it. - if createErr := c.Create(ctx, cm); createErr != nil { - return fmt.Errorf("creating migration status ConfigMap: %w", createErr) + // Reset on every write in case the Deployment was recreated with a new UID. + cm.OwnerReferences = []metav1.OwnerReference{ownerReference(deployment)} + cm.Data = map[string]string{ + "version": version, + "migratedAt": time.Now().UTC().Format(time.RFC3339), + "jobName": jobName, } return nil - } - - // Update existing ConfigMap (including OwnerReferences in case the Deployment - // was deleted and recreated with a new UID). - existing.Data = cm.Data - existing.Labels = cm.Labels - existing.OwnerReferences = cm.OwnerReferences - if updateErr := c.Update(ctx, existing); updateErr != nil { - return fmt.Errorf("updating migration status ConfigMap: %w", updateErr) - } - return nil -} - -// ensureDeploymentScaled ensures the Deployment is scaled to the desired replica count. -// The desired count is read from the AnnotationDesiredReplicas annotation. -// Returns true if the Deployment was already at the desired scale. -func ensureDeploymentScaled(ctx context.Context, c client.Client, deployment *appsv1.Deployment) (bool, error) { - desiredStr, ok := deployment.Annotations[AnnotationDesiredReplicas] - if !ok || desiredStr == "" { - // No annotation — nothing to do. The Deployment may not have been scaled down yet. - return true, nil - } - - desired, err := strconv.ParseInt(desiredStr, 10, 32) + }) if err != nil { - return false, fmt.Errorf("parsing desired replicas annotation: %w", err) + return fmt.Errorf("updating migration status ConfigMap: %w", err) } - - desiredInt32 := int32(desired) - current := int32(1) - if deployment.Spec.Replicas != nil { - current = *deployment.Spec.Replicas - } - - if current == desiredInt32 { - return true, nil - } - - patch := client.MergeFrom(deployment.DeepCopy()) - deployment.Spec.Replicas = ptr.To(desiredInt32) - if patchErr := c.Patch(ctx, deployment, patch); patchErr != nil { - return false, fmt.Errorf("scaling deployment to %d replicas: %w", desiredInt32, patchErr) - } - return false, nil + return nil } diff --git a/operator/internal/controller/migration_controller.go b/operator/internal/controller/migration_controller.go index 019cfa93..026ac66d 100644 --- a/operator/internal/controller/migration_controller.go +++ b/operator/internal/controller/migration_controller.go @@ -3,7 +3,6 @@ package controller import ( "context" "fmt" - "strings" "time" appsv1 "k8s.io/api/apps/v1" @@ -12,23 +11,25 @@ import ( apierrors "k8s.io/apimachinery/pkg/api/errors" metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" "k8s.io/apimachinery/pkg/types" + "k8s.io/utils/ptr" ctrl "sigs.k8s.io/controller-runtime" "sigs.k8s.io/controller-runtime/pkg/builder" "sigs.k8s.io/controller-runtime/pkg/client" - "sigs.k8s.io/controller-runtime/pkg/handler" "sigs.k8s.io/controller-runtime/pkg/log" "sigs.k8s.io/controller-runtime/pkg/predicate" - "sigs.k8s.io/controller-runtime/pkg/reconcile" ) -// MigrationReconciler watches OpenFGA Deployments and orchestrates database -// migrations when the application version changes. +// retryDelay is how long a failed migration Job is kept before it is replaced. +const retryDelay = 60 * time.Second + +// MigrationReconciler watches OpenFGA Deployments and runs a database +// migration Job whenever the OpenFGA image version changes. type MigrationReconciler struct { client.Client // BackoffLimit for migration Jobs. BackoffLimit int32 - // ActiveDeadlineSeconds for migration Jobs. + // ActiveDeadlineSeconds for migration Jobs; 0 means no deadline. ActiveDeadlineSeconds int64 // TTLSecondsAfterFinished for migration Jobs. TTLSecondsAfterFinished int32 @@ -38,226 +39,130 @@ type MigrationReconciler struct { func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) (ctrl.Result, error) { logger := log.FromContext(ctx) - // 1. Get the OpenFGA Deployment. deployment := &appsv1.Deployment{} if err := r.Get(ctx, req.NamespacedName, deployment); err != nil { - if apierrors.IsNotFound(err) { - return ctrl.Result{}, nil - } - return ctrl.Result{}, err + return ctrl.Result{}, client.IgnoreNotFound(err) } - - // 2. Skip if migration is not opted-in via annotation. - if len(deployment.Annotations) == 0 || deployment.Annotations[AnnotationMigrationEnabled] != "true" { - logger.V(1).Info("migration not enabled for this deployment, skipping") + if deployment.Annotations[AnnotationMigrationEnabled] != "true" { return ctrl.Result{}, nil } - // 3. Find the OpenFGA container and extract the desired version. - mainContainer, err := findOpenFGAContainer(deployment) + container, err := findOpenFGAContainer(deployment) if err != nil { - logger.Error(err, "unable to find OpenFGA container") return ctrl.Result{}, err } - desiredVersion := extractImageTag(mainContainer.Image) - - // 3b. Skip migration for memory datastore — just ensure the Deployment is scaled up. - if isMemoryDatastore(mainContainer) { - logger.V(1).Info("memory datastore detected, skipping migration") - if _, scaleErr := ensureDeploymentScaled(ctx, r.Client, deployment); scaleErr != nil { - return ctrl.Result{}, scaleErr - } - return ctrl.Result{}, nil - } - - // Migrations are keyed on the image reference. A mutable tag (e.g. "latest" - // or a floating "v1.14") can point at different images over time without the - // reference changing, so a rebuild that requires a migration will be missed. - if isMutableImageReference(mainContainer.Image) { - logger.Info("openfga image is not pinned to an immutable tag or digest; a migration may be silently skipped if the image changes without the tag changing", - "image", mainContainer.Image, "version", desiredVersion) - } - - // 4. Check current migration status from ConfigMap. - configMap := &corev1.ConfigMap{} - cmName := migrationConfigMapName(req.Name) - err = r.Get(ctx, types.NamespacedName{Name: cmName, Namespace: req.Namespace}, configMap) + desiredVersion := extractImageTag(container.Image) - currentVersion := "" - if err == nil { - currentVersion = configMap.Data["version"] - } else if !apierrors.IsNotFound(err) { + status := &corev1.ConfigMap{} + err = r.Get(ctx, types.NamespacedName{Name: migrationConfigMapName(req.Name), Namespace: req.Namespace}, status) + if err != nil && !apierrors.IsNotFound(err) { return ctrl.Result{}, fmt.Errorf("getting migration status: %w", err) } - - // 5. If versions match, ensure Deployment is scaled up and return. + currentVersion := status.Data["version"] if currentVersion == desiredVersion { - logger.V(1).Info("migration up to date", "version", desiredVersion) - // Strategic merge patches the condition by type without replacing the whole - // conditions list, so it won't clobber Available/Progressing. - statusPatch := client.StrategicMergeFrom(deployment.DeepCopy()) - if clearMigrationFailedCondition(deployment) { - if patchErr := r.Status().Patch(ctx, deployment, statusPatch); patchErr != nil { - return ctrl.Result{}, fmt.Errorf("clearing MigrationFailed condition: %w", patchErr) - } - } - if _, scaleErr := ensureDeploymentScaled(ctx, r.Client, deployment); scaleErr != nil { - return ctrl.Result{}, scaleErr - } - return ctrl.Result{}, nil - } - - logger.Info("migration needed", "currentVersion", currentVersion, "desiredVersion", desiredVersion) - - // 6. Check retry-after annotation to honor backoff cooldown. - if retryAfter, ok := deployment.Annotations[AnnotationRetryAfter]; ok { - retryTime, parseErr := time.Parse(time.RFC3339, retryAfter) - if parseErr == nil && time.Now().Before(retryTime) { - remaining := time.Until(retryTime) - logger.V(1).Info("in retry cooldown", "retryAfter", retryAfter, "remaining", remaining) - return ctrl.Result{RequeueAfter: remaining}, nil - } + _, err := r.patchCondition(ctx, deployment, clearMigrationFailedCondition) + return ctrl.Result{}, err } - // 7. Check if a migration Job already exists. - jobName := migrationJobName(req.Name) job := &batchv1.Job{} - err = r.Get(ctx, types.NamespacedName{Name: jobName, Namespace: req.Namespace}, job) - + err = r.Get(ctx, types.NamespacedName{Name: migrationJobName(req.Name), Namespace: req.Namespace}, job) if apierrors.IsNotFound(err) { - // Create the migration Job. - job = buildMigrationJob( - deployment, - mainContainer, - desiredVersion, - r.BackoffLimit, - r.ActiveDeadlineSeconds, - r.TTLSecondsAfterFinished, - ) - if createErr := r.Create(ctx, job); createErr != nil { - if apierrors.IsAlreadyExists(createErr) { - // A concurrent reconcile already created the Job; requeue to pick it up. - logger.V(1).Info("migration job already exists, will recheck", "job", jobName) + job = r.buildMigrationJob(deployment, container, desiredVersion) + if err := r.Create(ctx, job); err != nil { + if apierrors.IsAlreadyExists(err) { + // The cache has not caught up with a Job created by an earlier reconcile. return ctrl.Result{RequeueAfter: 5 * time.Second}, nil } - // Leave the retry-after annotation intact so the cooldown survives this failure. - return ctrl.Result{}, fmt.Errorf("creating migration job: %w", createErr) - } - // Clear the retry-after annotation now that the Job is created. - if _, hasRetry := deployment.Annotations[AnnotationRetryAfter]; hasRetry { - patch := client.MergeFrom(deployment.DeepCopy()) - delete(deployment.Annotations, AnnotationRetryAfter) - if patchErr := r.Patch(ctx, deployment, patch); patchErr != nil { - logger.Error(patchErr, "failed to clear retry-after annotation") - } + return ctrl.Result{}, fmt.Errorf("creating migration job: %w", err) } - logger.Info("created migration job", "job", jobName, "version", desiredVersion) - return ctrl.Result{RequeueAfter: 5 * time.Second}, nil - } else if err != nil { + logger.Info("created migration job", "job", job.Name, "currentVersion", currentVersion, "desiredVersion", desiredVersion) + return ctrl.Result{RequeueAfter: 10 * time.Second}, nil + } + if err != nil { return ctrl.Result{}, fmt.Errorf("getting migration job: %w", err) } - // 8. If the existing Job is for a different (or unknown) version, delete it - // and recreate. Check annotation first (supports digests > 63 chars), fall - // back to label. A Job with neither marker is treated as stale: we cannot - // trust its outcome to represent the current desired version, so trusting - // JobComplete in step 9 would write a wrong version into the status ConfigMap. + // Only a Job this operator created for the desired version is trusted. + // Anything else under the same name, such as a Job for a previous image or + // the chart's legacy Helm hook Job, is replaced. jobVersion := job.Annotations[AnnotationDesiredVersion] - versionMatch := jobVersion == desiredVersion - if jobVersion == "" { - // Label values have ":" replaced with "_", so sanitize desiredVersion for comparison. - sanitized := strings.ReplaceAll(desiredVersion, ":", "_") - if len(sanitized) > 63 { - sanitized = sanitized[:63] - } - jobVersion = job.Labels[LabelVersion] - versionMatch = jobVersion != "" && jobVersion == sanitized + complete := isJobConditionTrue(job, batchv1.JobComplete) + failedAt, failed := jobFailedAt(job) + outdated := jobVersion != desiredVersion + // A Job whose pod cannot start (a bad secret reference, an image pull + // error, an unschedulable pod) never fails on its own, so rebuild it once + // the Deployment's pod template has changed. A Job with a ready pod is left + // alone so a running migration is not cut off; one whose pod has just + // finished may still be rebuilt, which only re-runs a no-op migration. + if !outdated && !complete && !failed && job.Status.Active > 0 && ptr.Deref(job.Status.Ready, 0) == 0 { + want := r.buildMigrationJob(deployment, container, desiredVersion) + outdated = job.Annotations[AnnotationPodTemplateHash] != want.Annotations[AnnotationPodTemplateHash] } - if !versionMatch { - logger.Info("existing migration job is for a different or unknown version, deleting", "jobVersion", jobVersion, "desiredVersion", desiredVersion) - propagation := metav1.DeletePropagationBackground - if delErr := r.Delete(ctx, job, &client.DeleteOptions{ - PropagationPolicy: &propagation, - }); delErr != nil && !apierrors.IsNotFound(delErr) { - return ctrl.Result{}, fmt.Errorf("deleting stale migration job: %w", delErr) + if outdated { + logger.Info("replacing migration job", "job", job.Name, "jobVersion", jobVersion, "desiredVersion", desiredVersion) + if err := r.deleteJob(ctx, job); err != nil { + return ctrl.Result{}, err } return ctrl.Result{RequeueAfter: 5 * time.Second}, nil } - // 9. Check Job status using conditions for authoritative completion signals. - if isJobConditionTrue(job, batchv1.JobComplete) { - logger.Info("migration succeeded", "version", desiredVersion) - - // Clear MigrationFailed condition. - statusPatch := client.StrategicMergeFrom(deployment.DeepCopy()) - if clearMigrationFailedCondition(deployment) { - if patchErr := r.Status().Patch(ctx, deployment, statusPatch); patchErr != nil { - return ctrl.Result{}, fmt.Errorf("clearing MigrationFailed condition: %w", patchErr) - } - } - - // Update migration status ConfigMap. - if statusErr := updateMigrationStatus(ctx, r.Client, deployment, desiredVersion, jobName); statusErr != nil { - return ctrl.Result{}, statusErr - } - - // Scale Deployment back up. - if _, scaleErr := ensureDeploymentScaled(ctx, r.Client, deployment); scaleErr != nil { - return ctrl.Result{}, scaleErr + if complete { + if err := updateMigrationStatus(ctx, r.Client, deployment, desiredVersion, job.Name); err != nil { + return ctrl.Result{}, err } - - return ctrl.Result{}, nil + logger.Info("migration succeeded", "version", desiredVersion) + _, err := r.patchCondition(ctx, deployment, clearMigrationFailedCondition) + return ctrl.Result{}, err } - // JobFailureTarget is set as soon as the Job controller decides the Job - // will fail (backoff limit reached, deadline exceeded, etc.); JobFailed - // only flips after pods finish terminating, which can take BackoffLimit × - // ActiveDeadlineSeconds. Treating either as "failed" surfaces the failure - // to users within seconds instead of minutes. - if isJobConditionTrue(job, batchv1.JobFailed) || isJobConditionTrue(job, batchv1.JobFailureTarget) { - logger.Info("migration job failed, will delete and retry", "job", jobName, "version", desiredVersion) - - // Set MigrationFailed so kubectl describe shows the failure. - statusPatch := client.StrategicMergeFrom(deployment.DeepCopy()) - if setMigrationFailedCondition(deployment, desiredVersion) { - if patchErr := r.Status().Patch(ctx, deployment, statusPatch); patchErr != nil { - return ctrl.Result{}, fmt.Errorf("setting MigrationFailed condition: %w", patchErr) - } + if failed { + changed, err := r.patchCondition(ctx, deployment, func(d *appsv1.Deployment) bool { + return setMigrationFailedCondition(d, desiredVersion) + }) + if err != nil { + return ctrl.Result{}, err } - - // Persist a retry-after annotation so the cooldown is honored even - // when the Job deletion triggers an immediate re-enqueue. - retryAfter := time.Now().Add(60 * time.Second).UTC().Format(time.RFC3339) - patch := client.MergeFrom(deployment.DeepCopy()) - if deployment.Annotations == nil { - deployment.Annotations = make(map[string]string) + if changed { + logger.Info("migration job failed", "job", job.Name, "version", desiredVersion, "retryIn", retryDelay) } - deployment.Annotations[AnnotationRetryAfter] = retryAfter - if patchErr := r.Patch(ctx, deployment, patch); patchErr != nil { - return ctrl.Result{}, fmt.Errorf("persisting retry-after annotation: %w", patchErr) + // The failed Job itself is the retry timer, so the delay survives + // operator restarts and leaves the pod logs around to inspect. + if wait := retryDelay - time.Since(failedAt); wait > 0 { + return ctrl.Result{RequeueAfter: wait}, nil } - - // Delete the failed Job so a fresh one is created on the next reconcile. - propagation := metav1.DeletePropagationBackground - if delErr := r.Delete(ctx, job, &client.DeleteOptions{ - PropagationPolicy: &propagation, - }); delErr != nil && !apierrors.IsNotFound(delErr) { - return ctrl.Result{}, fmt.Errorf("deleting failed migration job: %w", delErr) + logger.Info("retrying migration", "job", job.Name, "version", desiredVersion) + if err := r.deleteJob(ctx, job); err != nil { + return ctrl.Result{}, err } - logger.Info("deleted failed migration job, will retry", "job", jobName) - - // Requeue after the cooldown period. - return ctrl.Result{RequeueAfter: 60 * time.Second}, nil + return ctrl.Result{RequeueAfter: 5 * time.Second}, nil } - // 10. Job still running — requeue. - logger.V(1).Info("migration job in progress", "job", jobName) return ctrl.Result{RequeueAfter: 10 * time.Second}, nil } -// isJobConditionTrue returns true if the Job has a condition of the given type -// with status True. This is more reliable than comparing status counters because -// the Job controller sets conditions atomically when it makes its final decision. +func (r *MigrationReconciler) deleteJob(ctx context.Context, job *batchv1.Job) error { + err := r.Delete(ctx, job, client.PropagationPolicy(metav1.DeletePropagationBackground)) + if client.IgnoreNotFound(err) != nil { + return fmt.Errorf("deleting migration job %s: %w", job.Name, err) + } + return nil +} + +// patchCondition applies update to the Deployment's status conditions and +// patches the status only if something changed. The strategic merge patch +// merges conditions by type, so the Deployment controller's own conditions are +// left alone. +func (r *MigrationReconciler) patchCondition(ctx context.Context, deployment *appsv1.Deployment, update func(*appsv1.Deployment) bool) (bool, error) { + patch := client.StrategicMergeFrom(deployment.DeepCopy()) + if !update(deployment) { + return false, nil + } + if err := r.Status().Patch(ctx, deployment, patch); err != nil { + return false, fmt.Errorf("patching MigrationFailed condition: %w", err) + } + return true, nil +} + func isJobConditionTrue(job *batchv1.Job, conditionType batchv1.JobConditionType) bool { for _, c := range job.Status.Conditions { if c.Type == conditionType && c.Status == corev1.ConditionTrue { @@ -267,28 +172,24 @@ func isJobConditionTrue(job *batchv1.Job, conditionType batchv1.JobConditionType return false } -// isMemoryDatastore checks if the Deployment is using the memory datastore -// (no database migration needed). -// -// NOTE: This only inspects explicit env vars on the container spec. If -// OPENFGA_DATASTORE_ENGINE is injected via envFrom (ConfigMap/Secret), it -// will not be detected here and the operator will attempt a migration. -func isMemoryDatastore(container *corev1.Container) bool { - for _, env := range container.Env { - if env.Name == "OPENFGA_DATASTORE_ENGINE" { - return strings.EqualFold(env.Value, "memory") +// jobFailedAt reports whether the Job has failed and when. JobFailureTarget is +// set as soon as the Job controller decides the Job will fail; JobFailed only +// once its pods have terminated. +func jobFailedAt(job *batchv1.Job) (time.Time, bool) { + for _, c := range job.Status.Conditions { + if (c.Type == batchv1.JobFailureTarget || c.Type == batchv1.JobFailed) && c.Status == corev1.ConditionTrue { + return c.LastTransitionTime.Time, true } } - return false + return time.Time{}, false } // setMigrationFailedCondition sets a MigrationFailed condition on the Deployment -// and reports whether anything actually changed, so callers can skip a no-op -// status write. LastTransitionTime only advances on a real status transition. +// and reports whether anything changed. LastTransitionTime only advances on a +// real status transition. // -// NOTE: this writes a custom condition type onto a built-in Deployment's status. -// meta.SetStatusCondition cannot be used here because Deployment.Status.Conditions -// is []appsv1.DeploymentCondition, not []metav1.Condition. +// Deployment.Status.Conditions is []appsv1.DeploymentCondition rather than +// []metav1.Condition, so meta.SetStatusCondition cannot be used. func setMigrationFailedCondition(deployment *appsv1.Deployment, version string) bool { message := fmt.Sprintf("Database migration failed for version %s. Check migration job logs.", version) for i, c := range deployment.Status.Conditions { @@ -316,10 +217,8 @@ func setMigrationFailedCondition(deployment *appsv1.Deployment, version string) } // clearMigrationFailedCondition sets an existing MigrationFailed condition to -// False and reports whether anything changed. When the condition is absent or -// already False it is a no-op — this is what stops the reconciler from patching -// status (and re-enqueueing the Deployment) on every reconcile of a healthy, -// version-matched Deployment. +// False and reports whether anything changed. It is a no-op when the condition +// is absent or already False, so healthy Deployments are never re-patched. func clearMigrationFailedCondition(deployment *appsv1.Deployment) bool { for i, c := range deployment.Status.Conditions { if c.Type == "MigrationFailed" { @@ -338,8 +237,7 @@ func clearMigrationFailedCondition(deployment *appsv1.Deployment) bool { // SetupWithManager sets up the controller with the Manager. func (r *MigrationReconciler) SetupWithManager(mgr ctrl.Manager) error { - // Only watch Deployments that are part of OpenFGA. - labelPredicate, err := predicate.LabelSelectorPredicate(metav1.LabelSelector{ + openfgaDeployments, err := predicate.LabelSelectorPredicate(metav1.LabelSelector{ MatchLabels: map[string]string{ LabelPartOf: LabelPartOfValue, LabelComponent: LabelComponentValue, @@ -349,29 +247,11 @@ func (r *MigrationReconciler) SetupWithManager(mgr ctrl.Manager) error { return fmt.Errorf("creating label predicate: %w", err) } + // Owning the status ConfigMap means deleting it triggers a reconcile, which + // runs the migration again once the previous Job is gone. return ctrl.NewControllerManagedBy(mgr). - For(&appsv1.Deployment{}, builder.WithPredicates(labelPredicate)). + For(&appsv1.Deployment{}, builder.WithPredicates(openfgaDeployments)). Owns(&batchv1.Job{}). - Watches(&corev1.ConfigMap{}, handler.EnqueueRequestsFromMapFunc( - func(ctx context.Context, obj client.Object) []reconcile.Request { - // Only watch ConfigMaps that are migration status ConfigMaps. - if obj.GetLabels()[LabelPartOf] != LabelPartOfValue || - obj.GetLabels()[LabelManagedBy] != LabelManagedByValue { - return nil - } - // Map back to the owning Deployment. - for _, ref := range obj.GetOwnerReferences() { - if ref.Kind == "Deployment" { - return []reconcile.Request{ - {NamespacedName: types.NamespacedName{ - Name: ref.Name, - Namespace: obj.GetNamespace(), - }}, - } - } - } - return nil - }, - )). + Owns(&corev1.ConfigMap{}). Complete(r) } diff --git a/operator/internal/controller/migration_controller_test.go b/operator/internal/controller/migration_controller_test.go index 7c744126..6d19ecaa 100644 --- a/operator/internal/controller/migration_controller_test.go +++ b/operator/internal/controller/migration_controller_test.go @@ -9,6 +9,7 @@ import ( appsv1 "k8s.io/api/apps/v1" batchv1 "k8s.io/api/batch/v1" corev1 "k8s.io/api/core/v1" + apierrors "k8s.io/apimachinery/pkg/api/errors" metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" "k8s.io/apimachinery/pkg/runtime" "k8s.io/apimachinery/pkg/types" @@ -20,17 +21,17 @@ import ( "sigs.k8s.io/controller-runtime/pkg/client/interceptor" ) -func newScheme() *runtime.Scheme { - s := runtime.NewScheme() - _ = clientgoscheme.AddToScheme(s) - return s -} +var ( + deploymentKey = types.NamespacedName{Name: "openfga", Namespace: "default"} + jobKey = types.NamespacedName{Name: "openfga-migrate", Namespace: "default"} + statusKey = types.NamespacedName{Name: "openfga-migration-status", Namespace: "default"} +) -func newTestDeployment(name, namespace, image string, replicas int32) *appsv1.Deployment { +func newTestDeployment(image string) *appsv1.Deployment { return &appsv1.Deployment{ ObjectMeta: metav1.ObjectMeta{ - Name: name, - Namespace: namespace, + Name: deploymentKey.Name, + Namespace: deploymentKey.Namespace, UID: "test-uid-123", Labels: map[string]string{ LabelPartOf: LabelPartOfValue, @@ -41,1193 +42,468 @@ func newTestDeployment(name, namespace, image string, replicas int32) *appsv1.De }, }, Spec: appsv1.DeploymentSpec{ - Replicas: ptr.To(replicas), - Selector: &metav1.LabelSelector{ - MatchLabels: map[string]string{"app": "openfga"}, - }, + Replicas: ptr.To(int32(3)), + Selector: &metav1.LabelSelector{MatchLabels: map[string]string{"app": "openfga"}}, Template: corev1.PodTemplateSpec{ - ObjectMeta: metav1.ObjectMeta{ - Labels: map[string]string{"app": "openfga"}, - }, + ObjectMeta: metav1.ObjectMeta{Labels: map[string]string{"app": "openfga"}}, Spec: corev1.PodSpec{ ServiceAccountName: "openfga", - Containers: []corev1.Container{ - { - Name: "openfga", - Image: image, - Env: []corev1.EnvVar{ - {Name: "OPENFGA_DATASTORE_ENGINE", Value: "postgres"}, - {Name: "OPENFGA_DATASTORE_URI", Value: "postgres://localhost/openfga"}, - {Name: "OPENFGA_LOG_LEVEL", Value: "info"}, - }, + Containers: []corev1.Container{{ + Name: "openfga", + Image: image, + Env: []corev1.EnvVar{ + {Name: "OPENFGA_DATASTORE_ENGINE", Value: "postgres"}, + {Name: "OPENFGA_DATASTORE_URI", Value: "postgres://localhost/openfga"}, + {Name: "OPENFGA_LOG_LEVEL", Value: "info"}, }, - }, + }}, }, }, }, } } -func newReconciler(objects ...runtime.Object) *MigrationReconciler { - scheme := newScheme() - clientBuilder := fake.NewClientBuilder().WithScheme(scheme). - WithStatusSubresource(&appsv1.Deployment{}) - for _, obj := range objects { - clientBuilder = clientBuilder.WithRuntimeObjects(obj) - } - return &MigrationReconciler{ - Client: clientBuilder.Build(), - BackoffLimit: DefaultBackoffLimit, - ActiveDeadlineSeconds: DefaultActiveDeadlineSeconds, - TTLSecondsAfterFinished: DefaultTTLSecondsAfterFinished, +// newTestJob returns the migration Job the operator builds for dep. +func newTestJob(dep *appsv1.Deployment, conditions ...batchv1.JobCondition) *batchv1.Job { + container := &dep.Spec.Template.Spec.Containers[0] + job := (&MigrationReconciler{}).buildMigrationJob(dep, container, extractImageTag(container.Image)) + job.Status.Conditions = conditions + return job +} + +func jobCondition(t batchv1.JobConditionType, at time.Time) batchv1.JobCondition { + return batchv1.JobCondition{Type: t, Status: corev1.ConditionTrue, LastTransitionTime: metav1.NewTime(at)} +} + +func newStatus(version string) *corev1.ConfigMap { + return &corev1.ConfigMap{ + ObjectMeta: metav1.ObjectMeta{Name: statusKey.Name, Namespace: statusKey.Namespace}, + Data: map[string]string{"version": version}, } } -func newReconcilerWithStatusPatchError(objects ...runtime.Object) *MigrationReconciler { - scheme := newScheme() - clientBuilder := fake.NewClientBuilder().WithScheme(scheme). - WithStatusSubresource(&appsv1.Deployment{}) - for _, obj := range objects { - clientBuilder = clientBuilder.WithRuntimeObjects(obj) - } - c := clientBuilder.WithInterceptorFuncs(interceptor.Funcs{ - SubResourcePatch: func( - ctx context.Context, - c client.Client, - subResourceName string, - obj client.Object, - patch client.Patch, - opts ...client.SubResourcePatchOption, - ) error { - if subResourceName == "status" { - return fmt.Errorf("simulated status patch error") - } - return c.SubResource(subResourceName).Patch(ctx, obj, patch, opts...) - }, - }).Build() +func newReconciler(t *testing.T, funcs *interceptor.Funcs, objects ...client.Object) *MigrationReconciler { + t.Helper() + scheme := runtime.NewScheme() + if err := clientgoscheme.AddToScheme(scheme); err != nil { + t.Fatal(err) + } + b := fake.NewClientBuilder().WithScheme(scheme). + WithStatusSubresource(&appsv1.Deployment{}). + WithObjects(objects...) + if funcs != nil { + b = b.WithInterceptorFuncs(*funcs) + } return &MigrationReconciler{ - Client: c, + Client: b.Build(), BackoffLimit: DefaultBackoffLimit, ActiveDeadlineSeconds: DefaultActiveDeadlineSeconds, TTLSecondsAfterFinished: DefaultTTLSecondsAfterFinished, } } -func findCondition(conditions []appsv1.DeploymentCondition, condType string) *appsv1.DeploymentCondition { - for i := range conditions { - if string(conditions[i].Type) == condType { - return &conditions[i] +var failStatusPatch = &interceptor.Funcs{ + SubResourcePatch: func(ctx context.Context, c client.Client, subResource string, obj client.Object, patch client.Patch, opts ...client.SubResourcePatchOption) error { + if subResource == "status" { + return fmt.Errorf("simulated status patch error") } - } - return nil + return c.SubResource(subResource).Patch(ctx, obj, patch, opts...) + }, } -func TestReconcile_FirstInstall_CreatesJob(t *testing.T) { - // Given: a Deployment with no migration-status ConfigMap. - dep := newTestDeployment("openfga", "default", "openfga/openfga:v1.14.0", 0) - r := newReconciler(dep) - - // When: reconciling. - result, err := r.Reconcile(context.Background(), ctrl.Request{ - NamespacedName: types.NamespacedName{Name: "openfga", Namespace: "default"}, - }) - - // Then: a migration Job should be created and requeue requested. +func reconcileOnce(t *testing.T, r *MigrationReconciler) ctrl.Result { + t.Helper() + result, err := r.Reconcile(context.Background(), ctrl.Request{NamespacedName: deploymentKey}) if err != nil { - t.Fatalf("unexpected error: %v", err) - } - if result.RequeueAfter == 0 { - t.Error("expected requeue, got none") - } - - // Verify the Job was created. - job := &batchv1.Job{} - if err := r.Get(context.Background(), types.NamespacedName{ - Name: "openfga-migrate", Namespace: "default", - }, job); err != nil { - t.Fatalf("expected migration job to be created: %v", err) - } - - if job.Spec.Template.Spec.Containers[0].Image != "openfga/openfga:v1.14.0" { - t.Errorf("expected job image openfga/openfga:v1.14.0, got %s", job.Spec.Template.Spec.Containers[0].Image) - } - - if job.Spec.Template.Spec.Containers[0].Args[0] != "migrate" { - t.Errorf("expected job args [migrate], got %v", job.Spec.Template.Spec.Containers[0].Args) - } - - // Verify all env vars from the main container were passed. - jobEnvNames := make(map[string]bool) - for _, env := range job.Spec.Template.Spec.Containers[0].Env { - jobEnvNames[env.Name] = true - } - for _, expected := range []string{"OPENFGA_DATASTORE_ENGINE", "OPENFGA_DATASTORE_URI", "OPENFGA_LOG_LEVEL"} { - if !jobEnvNames[expected] { - t.Errorf("expected env var %s to be passed to migration job", expected) - } + t.Fatalf("unexpected reconcile error: %v", err) } + return result } -func TestReconcile_FirstInstall_JobInheritsPullPolicyAndOwnerRef(t *testing.T) { - // Given: a Deployment whose OpenFGA container pins an explicit pull policy. - dep := newTestDeployment("openfga", "default", "openfga/openfga:v1.14.0", 0) - dep.Spec.Template.Spec.Containers[0].ImagePullPolicy = corev1.PullAlways - r := newReconciler(dep) - - // When: reconciling. - if _, err := r.Reconcile(context.Background(), ctrl.Request{ - NamespacedName: types.NamespacedName{Name: "openfga", Namespace: "default"}, - }); err != nil { - t.Fatalf("unexpected error: %v", err) +func getDeployment(t *testing.T, r *MigrationReconciler) *appsv1.Deployment { + t.Helper() + d := &appsv1.Deployment{} + if err := r.Get(context.Background(), deploymentKey, d); err != nil { + t.Fatalf("getting deployment: %v", err) } + return d +} +func getJob(r *MigrationReconciler) (*batchv1.Job, error) { job := &batchv1.Job{} - if err := r.Get(context.Background(), types.NamespacedName{ - Name: "openfga-migrate", Namespace: "default", - }, job); err != nil { - t.Fatalf("expected migration job to be created: %v", err) - } + return job, r.Get(context.Background(), jobKey, job) +} - // The migration Job must inherit the OpenFGA container's pull policy so it - // cannot run a stale cached image while the app pulls a fresh one. - if got := job.Spec.Template.Spec.Containers[0].ImagePullPolicy; got != corev1.PullAlways { - t.Errorf("expected job pull policy %q, got %q", corev1.PullAlways, got) - } +func getStatus(r *MigrationReconciler) (*corev1.ConfigMap, error) { + cm := &corev1.ConfigMap{} + return cm, r.Get(context.Background(), statusKey, cm) +} - // Owner reference makes the Job GC with the Deployment, but no blockOwnerDeletion: - // that needs the deployments/finalizers subresource the operator isn't granted, - // so the create would fail under OwnerReferencesPermissionEnforcement. - if len(job.OwnerReferences) != 1 { - t.Fatalf("expected exactly one owner reference, got %d", len(job.OwnerReferences)) - } - ref := job.OwnerReferences[0] - if ref.Controller == nil || !*ref.Controller { - t.Errorf("expected controller owner reference, got %+v", ref) - } - if ref.BlockOwnerDeletion != nil { - t.Errorf("expected blockOwnerDeletion to be unset, got %v", *ref.BlockOwnerDeletion) +func findCondition(conditions []appsv1.DeploymentCondition, condType string) *appsv1.DeploymentCondition { + for i := range conditions { + if string(conditions[i].Type) == condType { + return &conditions[i] + } } + return nil } -func TestReconcile_VersionMatch_ScalesUp(t *testing.T) { - // Given: a Deployment at 0 replicas with matching migration-status ConfigMap. - dep := newTestDeployment("openfga", "default", "openfga/openfga:v1.14.0", 0) - dep.Annotations[AnnotationDesiredReplicas] = "3" +func TestReconcile_FirstInstall_CreatesJob(t *testing.T) { + dep := newTestDeployment("openfga/openfga:v1.14.0") + dep.Annotations[AnnotationMigrationServiceAccount] = "openfga-migration" + dep.Spec.Template.Spec.Containers[0].ImagePullPolicy = corev1.PullAlways + r := newReconciler(t, nil, dep) - cm := &corev1.ConfigMap{ - ObjectMeta: metav1.ObjectMeta{ - Name: "openfga-migration-status", - Namespace: "default", - }, - Data: map[string]string{ - "version": "v1.14.0", - "migratedAt": "2026-04-06T12:00:00Z", - "jobName": "openfga-migrate", - }, + if result := reconcileOnce(t, r); result.RequeueAfter == 0 { + t.Error("expected a requeue to poll the new Job") } - r := newReconciler(dep, cm) - - // When: reconciling. - result, err := r.Reconcile(context.Background(), ctrl.Request{ - NamespacedName: types.NamespacedName{Name: "openfga", Namespace: "default"}, - }) - - // Then: no error, no requeue. + job, err := getJob(r) if err != nil { - t.Fatalf("unexpected error: %v", err) + t.Fatalf("expected migration job to be created: %v", err) } - if result.RequeueAfter != 0 { - t.Error("expected no requeue when versions match") + c := job.Spec.Template.Spec.Containers[0] + if c.Image != "openfga/openfga:v1.14.0" || len(c.Args) != 1 || c.Args[0] != "migrate" { + t.Errorf("unexpected migrate container: image=%s args=%v", c.Image, c.Args) } - - // Verify Deployment was scaled up. - updated := &appsv1.Deployment{} - if err := r.Get(context.Background(), types.NamespacedName{ - Name: "openfga", Namespace: "default", - }, updated); err != nil { - t.Fatalf("getting deployment: %v", err) + if c.ImagePullPolicy != corev1.PullAlways { + t.Errorf("expected the OpenFGA container's pull policy, got %q", c.ImagePullPolicy) } - if *updated.Spec.Replicas != 3 { - t.Errorf("expected 3 replicas, got %d", *updated.Spec.Replicas) + if len(c.Env) != 3 { + t.Errorf("expected all OpenFGA env vars to be passed through, got %v", c.Env) } -} - -func TestReconcile_VersionMatch_StatusPatchFailureStopsScaleUp(t *testing.T) { - dep := newTestDeployment("openfga", "default", "openfga/openfga:v1.14.0", 0) - dep.Annotations[AnnotationDesiredReplicas] = "3" - dep.Status.Conditions = []appsv1.DeploymentCondition{{ - Type: "MigrationFailed", - Status: corev1.ConditionTrue, - }} - cm := &corev1.ConfigMap{ - ObjectMeta: metav1.ObjectMeta{ - Name: "openfga-migration-status", - Namespace: "default", - }, - Data: map[string]string{"version": "v1.14.0"}, + if sa := job.Spec.Template.Spec.ServiceAccountName; sa != "openfga-migration" { + t.Errorf("expected migration service account, got %q", sa) } - r := newReconcilerWithStatusPatchError(dep, cm) - - if _, err := r.Reconcile(context.Background(), ctrl.Request{ - NamespacedName: types.NamespacedName{Name: "openfga", Namespace: "default"}, - }); err == nil { - t.Fatal("expected status patch error") + if got := job.Annotations[AnnotationDesiredVersion]; got != "v1.14.0" { + t.Errorf("expected desired-version annotation v1.14.0, got %q", got) } - - updated := &appsv1.Deployment{} - if err := r.Get(context.Background(), types.NamespacedName{ - Name: "openfga", Namespace: "default", - }, updated); err != nil { - t.Fatalf("getting deployment: %v", err) + if job.Spec.ActiveDeadlineSeconds != nil { + t.Errorf("expected no deadline by default, got %d", *job.Spec.ActiveDeadlineSeconds) } - if *updated.Spec.Replicas != 0 { - t.Errorf("expected replicas to remain at 0, got %d", *updated.Spec.Replicas) + if len(job.OwnerReferences) != 1 || !ptr.Deref(job.OwnerReferences[0].Controller, false) || job.OwnerReferences[0].BlockOwnerDeletion != nil { + t.Errorf("expected a single controller owner reference without blockOwnerDeletion, got %+v", job.OwnerReferences) } - cond := findCondition(updated.Status.Conditions, "MigrationFailed") - if cond == nil || cond.Status != corev1.ConditionTrue { - t.Fatalf("expected MigrationFailed condition to remain True, got %+v", cond) + if replicas := *getDeployment(t, r).Spec.Replicas; replicas != 3 { + t.Errorf("replicas must not be changed, got %d", replicas) } } -func TestReconcile_JobSucceeded_UpdatesConfigMapAndScalesUp(t *testing.T) { - // Given: a Deployment at 0 replicas, no ConfigMap, a succeeded migration Job, - // and a pre-existing MigrationFailed condition from a prior attempt. - dep := newTestDeployment("openfga", "default", "openfga/openfga:v1.14.0", 0) - dep.Annotations[AnnotationDesiredReplicas] = "3" - dep.Status.Conditions = []appsv1.DeploymentCondition{ - { - Type: "MigrationFailed", - Status: corev1.ConditionTrue, - Reason: "MigrationJobFailed", - Message: "Database migration failed for version v1.13.0.", - }, - } +func TestReconcile_FirstInstall_DeadlineAndDefaultServiceAccount(t *testing.T) { + r := newReconciler(t, nil, newTestDeployment("openfga/openfga:v1.14.0")) + r.ActiveDeadlineSeconds = 600 + reconcileOnce(t, r) - job := &batchv1.Job{ - ObjectMeta: metav1.ObjectMeta{ - Name: "openfga-migrate", - Namespace: "default", - Annotations: map[string]string{ - "openfga.dev/desired-version": "v1.14.0", - }, - OwnerReferences: []metav1.OwnerReference{ - { - APIVersion: "apps/v1", - Kind: "Deployment", - Name: "openfga", - UID: "test-uid-123", - }, - }, - }, - Spec: batchv1.JobSpec{ - BackoffLimit: ptr.To(int32(3)), - Template: corev1.PodTemplateSpec{ - Spec: corev1.PodSpec{ - Containers: []corev1.Container{{Name: "migrate", Image: "openfga/openfga:v1.14.0"}}, - RestartPolicy: corev1.RestartPolicyNever, - }, - }, - }, - Status: batchv1.JobStatus{ - Succeeded: 1, - Conditions: []batchv1.JobCondition{ - { - Type: batchv1.JobComplete, - Status: corev1.ConditionTrue, - }, - }, - }, - } - - r := newReconciler(dep, job) - - // When: reconciling. - _, err := r.Reconcile(context.Background(), ctrl.Request{ - NamespacedName: types.NamespacedName{Name: "openfga", Namespace: "default"}, - }) - - // Then: no error. + job, err := getJob(r) if err != nil { - t.Fatalf("unexpected error: %v", err) - } - - // Verify ConfigMap was created. - cm := &corev1.ConfigMap{} - if err := r.Get(context.Background(), types.NamespacedName{ - Name: "openfga-migration-status", Namespace: "default", - }, cm); err != nil { - t.Fatalf("expected ConfigMap to be created: %v", err) - } - if cm.Data["version"] != "v1.14.0" { - t.Errorf("expected version v1.14.0 in ConfigMap, got %s", cm.Data["version"]) - } - - // Verify Deployment was scaled up. - updated := &appsv1.Deployment{} - if err := r.Get(context.Background(), types.NamespacedName{ - Name: "openfga", Namespace: "default", - }, updated); err != nil { - t.Fatalf("getting deployment: %v", err) - } - if *updated.Spec.Replicas != 3 { - t.Errorf("expected 3 replicas, got %d", *updated.Spec.Replicas) - } - - // Verify MigrationFailed condition was cleared. - cond := findCondition(updated.Status.Conditions, "MigrationFailed") - if cond == nil { - t.Fatal("expected MigrationFailed condition to exist") + t.Fatalf("expected migration job to be created: %v", err) } - if cond.Status != corev1.ConditionFalse { - t.Errorf("expected MigrationFailed status False after success, got %s", cond.Status) + if got := ptr.Deref(job.Spec.ActiveDeadlineSeconds, 0); got != 600 { + t.Errorf("expected activeDeadlineSeconds 600, got %d", got) } - if cond.Reason != "MigrationSucceeded" { - t.Errorf("expected reason MigrationSucceeded, got %s", cond.Reason) + if sa := job.Spec.Template.Spec.ServiceAccountName; sa != "openfga" { + t.Errorf("expected the Deployment's service account, got %q", sa) } } -func TestReconcile_JobFailed_SetsRetryAnnotationAndRequeues(t *testing.T) { - // Given: a Deployment at 0 replicas and a failed migration Job. - dep := newTestDeployment("openfga", "default", "openfga/openfga:v1.14.0", 0) - dep.Annotations[AnnotationDesiredReplicas] = "3" - - job := &batchv1.Job{ - ObjectMeta: metav1.ObjectMeta{ - Name: "openfga-migrate", - Namespace: "default", - Annotations: map[string]string{ - "openfga.dev/desired-version": "v1.14.0", - }, - OwnerReferences: []metav1.OwnerReference{ - { - APIVersion: "apps/v1", - Kind: "Deployment", - Name: "openfga", - UID: "test-uid-123", - }, - }, - }, - Spec: batchv1.JobSpec{ - BackoffLimit: ptr.To(int32(3)), - Template: corev1.PodTemplateSpec{ - Spec: corev1.PodSpec{ - Containers: []corev1.Container{{Name: "migrate", Image: "openfga/openfga:v1.14.0"}}, - RestartPolicy: corev1.RestartPolicyNever, - }, - }, +func TestReconcile_JobAlreadyExistsOnCreate_Requeues(t *testing.T) { + r := newReconciler(t, &interceptor.Funcs{ + Create: func(ctx context.Context, c client.WithWatch, obj client.Object, opts ...client.CreateOption) error { + return apierrors.NewAlreadyExists(batchv1.Resource("jobs"), obj.GetName()) }, - Status: batchv1.JobStatus{ - Failed: 3, - Conditions: []batchv1.JobCondition{ - { - Type: batchv1.JobFailed, - Status: corev1.ConditionTrue, - }, - }, - }, - } - - r := newReconciler(dep, job) + }, newTestDeployment("openfga/openfga:v1.14.0")) - // When: reconciling. - result, err := r.Reconcile(context.Background(), ctrl.Request{ - NamespacedName: types.NamespacedName{Name: "openfga", Namespace: "default"}, - }) - - // Then: no error, but requeue after 60s for retry. - if err != nil { - t.Fatalf("unexpected error: %v", err) - } - if result.RequeueAfter != 60*time.Second { - t.Errorf("expected 60s requeue, got %v", result.RequeueAfter) - } - - // Verify Deployment replicas unchanged (still at 0 from fresh install). - updated := &appsv1.Deployment{} - if getErr := r.Get(context.Background(), types.NamespacedName{ - Name: "openfga", Namespace: "default", - }, updated); getErr != nil { - t.Fatalf("getting deployment: %v", getErr) - } - if *updated.Spec.Replicas != 0 { - t.Errorf("expected 0 replicas after failed migration, got %d", *updated.Spec.Replicas) + if result := reconcileOnce(t, r); result.RequeueAfter == 0 { + t.Error("expected a requeue when the Job already exists") } +} - // Verify the failed Job was deleted. - deletedJob := &batchv1.Job{} - if getErr := r.Get(context.Background(), types.NamespacedName{ - Name: "openfga-migrate", Namespace: "default", - }, deletedJob); getErr == nil { - t.Error("expected failed migration job to be deleted") - } +func TestReconcile_VersionMatch_ClearsFailedCondition(t *testing.T) { + dep := newTestDeployment("openfga/openfga:v1.14.0") + dep.Status.Conditions = []appsv1.DeploymentCondition{{Type: "MigrationFailed", Status: corev1.ConditionTrue}} + r := newReconciler(t, nil, dep, newStatus("v1.14.0")) - // Verify retry-after annotation was set on the Deployment. - if _, ok := updated.Annotations[AnnotationRetryAfter]; !ok { - t.Error("expected retry-after annotation to be set on Deployment") + if result := reconcileOnce(t, r); result.RequeueAfter != 0 { + t.Errorf("expected no requeue when versions match, got %v", result.RequeueAfter) } - - // Verify MigrationFailed condition was set. - cond := findCondition(updated.Status.Conditions, "MigrationFailed") - if cond == nil { - t.Fatal("expected MigrationFailed condition to be set") + if _, err := getJob(r); !apierrors.IsNotFound(err) { + t.Errorf("expected no migration job, got err=%v", err) } - if cond.Status != corev1.ConditionTrue { - t.Errorf("expected MigrationFailed status True, got %s", cond.Status) + updated := getDeployment(t, r) + if cond := findCondition(updated.Status.Conditions, "MigrationFailed"); cond == nil || cond.Status != corev1.ConditionFalse { + t.Errorf("expected MigrationFailed=False, got %+v", cond) } - if cond.Reason != "MigrationJobFailed" { - t.Errorf("expected reason MigrationJobFailed, got %s", cond.Reason) + if *updated.Spec.Replicas != 3 { + t.Errorf("replicas must not be changed, got %d", *updated.Spec.Replicas) } } -func TestReconcile_JobFailed_StatusPatchFailurePreservesJob(t *testing.T) { - dep := newTestDeployment("openfga", "default", "openfga/openfga:v1.14.0", 0) - dep.Annotations[AnnotationDesiredReplicas] = "3" - job := &batchv1.Job{ - ObjectMeta: metav1.ObjectMeta{ - Name: "openfga-migrate", - Namespace: "default", - Annotations: map[string]string{ - AnnotationDesiredVersion: "v1.14.0", - }, - }, - Status: batchv1.JobStatus{ - Conditions: []batchv1.JobCondition{{ - Type: batchv1.JobFailed, - Status: corev1.ConditionTrue, - }}, - }, - } - r := newReconcilerWithStatusPatchError(dep, job) - - if _, err := r.Reconcile(context.Background(), ctrl.Request{ - NamespacedName: types.NamespacedName{Name: "openfga", Namespace: "default"}, - }); err == nil { - t.Fatal("expected status patch error") - } +func TestReconcile_VersionMatch_StatusPatchError(t *testing.T) { + dep := newTestDeployment("openfga/openfga:v1.14.0") + dep.Status.Conditions = []appsv1.DeploymentCondition{{Type: "MigrationFailed", Status: corev1.ConditionTrue}} + r := newReconciler(t, failStatusPatch, dep, newStatus("v1.14.0")) - preservedJob := &batchv1.Job{} - if err := r.Get(context.Background(), types.NamespacedName{ - Name: "openfga-migrate", Namespace: "default", - }, preservedJob); err != nil { - t.Fatalf("expected failed Job to remain for retry: %v", err) - } - updated := &appsv1.Deployment{} - if err := r.Get(context.Background(), types.NamespacedName{ - Name: "openfga", Namespace: "default", - }, updated); err != nil { - t.Fatalf("getting deployment: %v", err) - } - if _, ok := updated.Annotations[AnnotationRetryAfter]; ok { - t.Error("retry-after must not be set when the failure condition was not persisted") + if _, err := r.Reconcile(context.Background(), ctrl.Request{NamespacedName: deploymentKey}); err == nil { + t.Fatal("expected the status patch error to be returned") } } -func TestReconcile_JobFailureTarget_TreatedAsFailed(t *testing.T) { - // Given: a Job with only JobFailureTarget=True (no JobFailed yet). The - // Job controller sets this as soon as it decides the Job will fail, - // before pods finish terminating and JobFailed is recorded. The operator - // should treat this as a failure to surface the error in seconds rather - // than waiting the full BackoffLimit × ActiveDeadlineSeconds. - dep := newTestDeployment("openfga", "default", "openfga/openfga:v1.14.0", 0) - dep.Annotations[AnnotationDesiredReplicas] = "3" +func TestReconcile_JobSucceeded_CreatesStatus(t *testing.T) { + dep := newTestDeployment("openfga/openfga:v1.14.0") + dep.Status.Conditions = []appsv1.DeploymentCondition{{Type: "MigrationFailed", Status: corev1.ConditionTrue, Reason: "MigrationJobFailed"}} + r := newReconciler(t, nil, dep, newTestJob(dep, jobCondition(batchv1.JobComplete, time.Now()))) - job := &batchv1.Job{ - ObjectMeta: metav1.ObjectMeta{ - Name: "openfga-migrate", - Namespace: "default", - Annotations: map[string]string{ - "openfga.dev/desired-version": "v1.14.0", - }, - OwnerReferences: []metav1.OwnerReference{ - { - APIVersion: "apps/v1", - Kind: "Deployment", - Name: "openfga", - UID: "test-uid-123", - }, - }, - }, - Status: batchv1.JobStatus{ - Conditions: []batchv1.JobCondition{ - {Type: batchv1.JobFailureTarget, Status: corev1.ConditionTrue, Reason: "BackoffLimitExceeded"}, - }, - }, + if result := reconcileOnce(t, r); result.RequeueAfter != 0 { + t.Errorf("expected no requeue after success, got %v", result.RequeueAfter) } - - r := newReconciler(dep, job) - - result, err := r.Reconcile(context.Background(), ctrl.Request{ - NamespacedName: types.NamespacedName{Name: "openfga", Namespace: "default"}, - }) + cm, err := getStatus(r) if err != nil { - t.Fatalf("unexpected error: %v", err) + t.Fatalf("expected migration status ConfigMap: %v", err) } - if result.RequeueAfter != 60*time.Second { - t.Errorf("expected 60s requeue, got %v", result.RequeueAfter) + if cm.Data["version"] != "v1.14.0" || cm.Data["jobName"] != jobKey.Name { + t.Errorf("unexpected status data: %v", cm.Data) } - - updated := &appsv1.Deployment{} - if getErr := r.Get(context.Background(), types.NamespacedName{ - Name: "openfga", Namespace: "default", - }, updated); getErr != nil { - t.Fatalf("getting deployment: %v", getErr) + if len(cm.OwnerReferences) != 1 || cm.OwnerReferences[0].UID != "test-uid-123" { + t.Errorf("expected the Deployment to own the status ConfigMap, got %+v", cm.OwnerReferences) } - - deletedJob := &batchv1.Job{} - if getErr := r.Get(context.Background(), types.NamespacedName{ - Name: "openfga-migrate", Namespace: "default", - }, deletedJob); getErr == nil { - t.Error("expected migration job to be deleted on JobFailureTarget") - } - if _, ok := updated.Annotations[AnnotationRetryAfter]; !ok { - t.Error("expected retry-after annotation to be set") - } - cond := findCondition(updated.Status.Conditions, "MigrationFailed") - if cond == nil || cond.Status != corev1.ConditionTrue { - t.Fatal("expected MigrationFailed condition True") + cond := findCondition(getDeployment(t, r).Status.Conditions, "MigrationFailed") + if cond == nil || cond.Status != corev1.ConditionFalse || cond.Reason != "MigrationSucceeded" { + t.Errorf("expected MigrationFailed=False/MigrationSucceeded, got %+v", cond) } } -func TestReconcile_RetryAfterCooldown_SkipsJobCreation(t *testing.T) { - // Given: a Deployment with a retry-after annotation in the future. - dep := newTestDeployment("openfga", "default", "openfga/openfga:v1.14.0", 0) - dep.Annotations[AnnotationDesiredReplicas] = "3" - dep.Annotations[AnnotationRetryAfter] = time.Now().Add(30 * time.Second).UTC().Format(time.RFC3339) - - r := newReconciler(dep) +func TestReconcile_JobSucceeded_UpdatesStatus(t *testing.T) { + status := newStatus("v1.13.0") + status.OwnerReferences = []metav1.OwnerReference{{APIVersion: "apps/v1", Kind: "Deployment", Name: "openfga", UID: "old-uid"}} + dep := newTestDeployment("openfga/openfga:v1.14.0") + r := newReconciler(t, nil, dep, status, newTestJob(dep, jobCondition(batchv1.JobComplete, time.Now()))) - // When: reconciling. - result, err := r.Reconcile(context.Background(), ctrl.Request{ - NamespacedName: types.NamespacedName{Name: "openfga", Namespace: "default"}, - }) + reconcileOnce(t, r) - // Then: no error, requeue with remaining cooldown time. + cm, err := getStatus(r) if err != nil { - t.Fatalf("unexpected error: %v", err) - } - if result.RequeueAfter == 0 { - t.Error("expected requeue during cooldown") + t.Fatalf("expected migration status ConfigMap: %v", err) } - if result.RequeueAfter > 30*time.Second { - t.Errorf("expected requeue within 30s, got %v", result.RequeueAfter) + if cm.Data["version"] != "v1.14.0" { + t.Errorf("expected version v1.14.0, got %q", cm.Data["version"]) } - - // Verify no Job was created. - job := &batchv1.Job{} - if getErr := r.Get(context.Background(), types.NamespacedName{ - Name: "openfga-migrate", Namespace: "default", - }, job); getErr == nil { - t.Error("expected no migration job during cooldown") + if cm.OwnerReferences[0].UID != "test-uid-123" { + t.Errorf("expected owner reference to be reset to the current Deployment, got %+v", cm.OwnerReferences) } } -func TestReconcile_UnknownVersionJob_DeletedNotTrusted(t *testing.T) { - // Given: a Deployment desiring v1.14.0 and a JobComplete migration Job that - // carries no version annotation or label (e.g. left over from an older - // operator or created by a third-party tool). Trusting its outcome would - // write the wrong version into the migration-status ConfigMap. - dep := newTestDeployment("openfga", "default", "openfga/openfga:v1.14.0", 0) - dep.Annotations[AnnotationDesiredReplicas] = "3" - - job := &batchv1.Job{ - ObjectMeta: metav1.ObjectMeta{ - Name: "openfga-migrate", - Namespace: "default", - OwnerReferences: []metav1.OwnerReference{ - { - APIVersion: "apps/v1", - Kind: "Deployment", - Name: "openfga", - UID: "test-uid-123", - }, - }, - }, - Status: batchv1.JobStatus{ - Conditions: []batchv1.JobCondition{ - {Type: batchv1.JobComplete, Status: corev1.ConditionTrue}, - }, - }, - } - - r := newReconciler(dep, job) - - // When: reconciling. - result, err := r.Reconcile(context.Background(), ctrl.Request{ - NamespacedName: types.NamespacedName{Name: "openfga", Namespace: "default"}, - }) - - // Then: the Job is deleted and a requeue is scheduled; the ConfigMap is - // NOT created from the unknown-version Job's outcome. - if err != nil { - t.Fatalf("unexpected error: %v", err) - } - if result.RequeueAfter == 0 { - t.Error("expected requeue after deleting unknown-version job") - } +func TestReconcile_JobInProgress_Requeues(t *testing.T) { + dep := newTestDeployment("openfga/openfga:v1.14.0") + r := newReconciler(t, nil, dep, newTestJob(dep)) - deletedJob := &batchv1.Job{} - if getErr := r.Get(context.Background(), types.NamespacedName{ - Name: "openfga-migrate", Namespace: "default", - }, deletedJob); getErr == nil { - t.Error("expected unknown-version job to be deleted") + if result := reconcileOnce(t, r); result.RequeueAfter != 10*time.Second { + t.Errorf("expected 10s requeue for an in-progress job, got %v", result.RequeueAfter) } - - cm := &corev1.ConfigMap{} - if getErr := r.Get(context.Background(), types.NamespacedName{ - Name: "openfga-migration-status", Namespace: "default", - }, cm); getErr == nil { - t.Errorf("expected no migration-status ConfigMap; got version=%q", cm.Data["version"]) + if _, err := getStatus(r); !apierrors.IsNotFound(err) { + t.Errorf("expected no status ConfigMap while the job runs, got err=%v", err) } } -func TestReconcile_RetryAfterPersistsOnJobCreateFailure(t *testing.T) { - // Given: a Deployment with an elapsed retry-after annotation, and a client - // that fails Job creation with a non-AlreadyExists error. - dep := newTestDeployment("openfga", "default", "openfga/openfga:v1.14.0", 0) - dep.Annotations[AnnotationDesiredReplicas] = "3" - dep.Annotations[AnnotationRetryAfter] = time.Now().Add(-1 * time.Second).UTC().Format(time.RFC3339) - - scheme := newScheme() - c := fake.NewClientBuilder(). - WithScheme(scheme). - WithStatusSubresource(&appsv1.Deployment{}). - WithRuntimeObjects(dep). - WithInterceptorFuncs(interceptor.Funcs{ - Create: func(ctx context.Context, c client.WithWatch, obj client.Object, opts ...client.CreateOption) error { - if _, ok := obj.(*batchv1.Job); ok { - return fmt.Errorf("simulated transient API error") - } - return c.Create(ctx, obj, opts...) - }, - }). - Build() - r := &MigrationReconciler{ - Client: c, - BackoffLimit: DefaultBackoffLimit, - ActiveDeadlineSeconds: DefaultActiveDeadlineSeconds, - TTLSecondsAfterFinished: DefaultTTLSecondsAfterFinished, - } - - // When: reconciling. - _, err := r.Reconcile(context.Background(), ctrl.Request{ - NamespacedName: types.NamespacedName{Name: "openfga", Namespace: "default"}, - }) +func TestReconcile_JobFailed_KeepsJobUntilRetryDelay(t *testing.T) { + dep := newTestDeployment("openfga/openfga:v1.14.0") + r := newReconciler(t, nil, dep, newTestJob(dep, jobCondition(batchv1.JobFailed, time.Now().Add(-10*time.Second)))) - // Then: an error is returned and the retry-after annotation is preserved - // so the next reconcile honors the cooldown. - if err == nil { - t.Fatal("expected error from failed job creation") + result := reconcileOnce(t, r) + if result.RequeueAfter <= 0 || result.RequeueAfter > retryDelay-10*time.Second { + t.Errorf("expected a requeue for the rest of the retry delay, got %v", result.RequeueAfter) } - - updated := &appsv1.Deployment{} - if getErr := r.Get(context.Background(), types.NamespacedName{ - Name: "openfga", Namespace: "default", - }, updated); getErr != nil { - t.Fatalf("getting deployment: %v", getErr) + if _, err := getJob(r); err != nil { + t.Errorf("expected the failed job to be kept during the retry delay: %v", err) } - if _, ok := updated.Annotations[AnnotationRetryAfter]; !ok { - t.Error("expected retry-after annotation to persist after Job creation failure") + cond := findCondition(getDeployment(t, r).Status.Conditions, "MigrationFailed") + if cond == nil || cond.Status != corev1.ConditionTrue || cond.Reason != "MigrationJobFailed" { + t.Errorf("expected MigrationFailed=True, got %+v", cond) } } -func TestReconcile_RetryAfterClearedAfterJobCreated(t *testing.T) { - // Given: a Deployment with an elapsed retry-after annotation. - dep := newTestDeployment("openfga", "default", "openfga/openfga:v1.14.0", 0) - dep.Annotations[AnnotationDesiredReplicas] = "3" - dep.Annotations[AnnotationRetryAfter] = time.Now().Add(-1 * time.Second).UTC().Format(time.RFC3339) +func TestReconcile_JobFailed_RetriesAfterDelay(t *testing.T) { + for _, condType := range []batchv1.JobConditionType{batchv1.JobFailed, batchv1.JobFailureTarget} { + t.Run(string(condType), func(t *testing.T) { + dep := newTestDeployment("openfga/openfga:v1.14.0") + r := newReconciler(t, nil, dep, newTestJob(dep, jobCondition(condType, time.Now().Add(-2*retryDelay)))) - r := newReconciler(dep) + reconcileOnce(t, r) + if _, err := getJob(r); !apierrors.IsNotFound(err) { + t.Fatalf("expected the failed job to be deleted, got err=%v", err) + } + if cond := findCondition(getDeployment(t, r).Status.Conditions, "MigrationFailed"); cond == nil || cond.Status != corev1.ConditionTrue { + t.Errorf("expected MigrationFailed=True, got %+v", cond) + } - // When: reconciling. - if _, err := r.Reconcile(context.Background(), ctrl.Request{ - NamespacedName: types.NamespacedName{Name: "openfga", Namespace: "default"}, - }); err != nil { - t.Fatalf("unexpected error: %v", err) + reconcileOnce(t, r) + job, err := getJob(r) + if err != nil { + t.Fatalf("expected a new migration job: %v", err) + } + if len(job.Status.Conditions) != 0 { + t.Errorf("expected a fresh job, got conditions %+v", job.Status.Conditions) + } + }) } +} - // Then: the Job exists and the retry-after annotation has been cleared. - job := &batchv1.Job{} - if getErr := r.Get(context.Background(), types.NamespacedName{ - Name: "openfga-migrate", Namespace: "default", - }, job); getErr != nil { - t.Fatalf("expected migration job to be created: %v", getErr) - } +func TestReconcile_JobFailed_StatusPatchErrorKeepsJob(t *testing.T) { + dep := newTestDeployment("openfga/openfga:v1.14.0") + r := newReconciler(t, failStatusPatch, dep, newTestJob(dep, jobCondition(batchv1.JobFailed, time.Now().Add(-2*retryDelay)))) - updated := &appsv1.Deployment{} - if getErr := r.Get(context.Background(), types.NamespacedName{ - Name: "openfga", Namespace: "default", - }, updated); getErr != nil { - t.Fatalf("getting deployment: %v", getErr) + if _, err := r.Reconcile(context.Background(), ctrl.Request{NamespacedName: deploymentKey}); err == nil { + t.Fatal("expected the status patch error to be returned") } - if _, ok := updated.Annotations[AnnotationRetryAfter]; ok { - t.Error("expected retry-after annotation to be cleared after Job created") + if _, err := getJob(r); err != nil { + t.Errorf("the failed job must be kept until the failure is recorded: %v", err) } } -func TestReconcile_MemoryDatastore_SkipsMigration(t *testing.T) { - // Given: a Deployment using the memory datastore. - dep := newTestDeployment("openfga", "default", "openfga/openfga:v1.14.0", 0) - dep.Annotations[AnnotationDesiredReplicas] = "1" - dep.Spec.Template.Spec.Containers[0].Env = []corev1.EnvVar{ - {Name: "OPENFGA_DATASTORE_ENGINE", Value: "memory"}, - } - - r := newReconciler(dep) - - // When: reconciling. - result, err := r.Reconcile(context.Background(), ctrl.Request{ - NamespacedName: types.NamespacedName{Name: "openfga", Namespace: "default"}, - }) - - // Then: no error, no requeue. - if err != nil { - t.Fatalf("unexpected error: %v", err) - } - if result.RequeueAfter != 0 { - t.Error("expected no requeue for memory datastore") - } +func TestReconcile_JobForOtherVersion_Replaced(t *testing.T) { + r := newReconciler(t, nil, newTestDeployment("openfga/openfga:v1.15.0"), + newTestJob(newTestDeployment("openfga/openfga:v1.14.0"), jobCondition(batchv1.JobComplete, time.Now()))) - // Verify Deployment was scaled up (no migration needed). - updated := &appsv1.Deployment{} - if getErr := r.Get(context.Background(), types.NamespacedName{ - Name: "openfga", Namespace: "default", - }, updated); getErr != nil { - t.Fatalf("getting deployment: %v", getErr) + if result := reconcileOnce(t, r); result.RequeueAfter == 0 { + t.Error("expected a requeue after deleting the stale job") } - if *updated.Spec.Replicas != 1 { - t.Errorf("expected 1 replica, got %d", *updated.Spec.Replicas) + if _, err := getJob(r); !apierrors.IsNotFound(err) { + t.Errorf("expected the stale job to be deleted, got err=%v", err) } - - // Verify no Job was created. - job := &batchv1.Job{} - if getErr := r.Get(context.Background(), types.NamespacedName{ - Name: "openfga-migrate", Namespace: "default", - }, job); getErr == nil { - t.Error("expected no migration job for memory datastore") + if _, err := getStatus(r); !apierrors.IsNotFound(err) { + t.Errorf("a job for another version must not be recorded as migrated, got err=%v", err) } } -func TestReconcile_DeploymentNotFound_NoError(t *testing.T) { - r := newReconciler() - - result, err := r.Reconcile(context.Background(), ctrl.Request{ - NamespacedName: types.NamespacedName{Name: "nonexistent", Namespace: "default"}, - }) +func TestReconcile_PendingJobWithOutdatedTemplate_Replaced(t *testing.T) { + dep := newTestDeployment("openfga/openfga:v1.14.0") + job := newTestJob(dep) + job.Status.Active = 1 + job.Status.Ready = ptr.To(int32(0)) + dep.Spec.Template.Spec.Containers[0].Env[1].Value = "postgres://db.example.com/openfga" + r := newReconciler(t, nil, dep, job) - if err != nil { - t.Fatalf("unexpected error: %v", err) - } - if result.RequeueAfter != 0 { - t.Error("expected no requeue for missing deployment") + reconcileOnce(t, r) + if _, err := getJob(r); !apierrors.IsNotFound(err) { + t.Fatalf("expected the job built from the old pod template to be deleted, got err=%v", err) } -} -func TestReconcile_FindContainerByName(t *testing.T) { - // Given: a Deployment with a sidecar before the openfga container. - dep := newTestDeployment("openfga", "default", "openfga/openfga:v1.14.0", 0) - dep.Spec.Template.Spec.Containers = []corev1.Container{ - { - Name: "sidecar", - Image: "envoyproxy/envoy:v1.30", - }, - { - Name: "openfga", - Image: "openfga/openfga:v1.14.0", - Env: []corev1.EnvVar{ - {Name: "OPENFGA_DATASTORE_ENGINE", Value: "postgres"}, - {Name: "OPENFGA_DATASTORE_URI", Value: "postgres://localhost/openfga"}, - }, - }, - } - - r := newReconciler(dep) - - // When: reconciling. - result, err := r.Reconcile(context.Background(), ctrl.Request{ - NamespacedName: types.NamespacedName{Name: "openfga", Namespace: "default"}, - }) - - // Then: Job should use the openfga container's image, not the sidecar's. + reconcileOnce(t, r) + rebuilt, err := getJob(r) if err != nil { - t.Fatalf("unexpected error: %v", err) + t.Fatalf("expected a new migration job: %v", err) } - if result.RequeueAfter == 0 { - t.Error("expected requeue, got none") - } - - job := &batchv1.Job{} - if err := r.Get(context.Background(), types.NamespacedName{ - Name: "openfga-migrate", Namespace: "default", - }, job); err != nil { - t.Fatalf("expected migration job to be created: %v", err) + if uri := rebuilt.Spec.Template.Spec.Containers[0].Env[1].Value; uri != "postgres://db.example.com/openfga" { + t.Errorf("expected the new job to use the updated env, got %q", uri) } +} - if job.Spec.Template.Spec.Containers[0].Image != "openfga/openfga:v1.14.0" { - t.Errorf("expected job image openfga/openfga:v1.14.0, got %s", job.Spec.Template.Spec.Containers[0].Image) +func TestReconcile_StartedJobWithOutdatedTemplate_Kept(t *testing.T) { + for _, tt := range []struct { + name string + status batchv1.JobStatus + }{ + {"pod running", batchv1.JobStatus{Active: 1, Ready: ptr.To(int32(1))}}, + {"pod finished before the job is marked complete", batchv1.JobStatus{Succeeded: 1, Ready: ptr.To(int32(0))}}, + } { + t.Run(tt.name, func(t *testing.T) { + dep := newTestDeployment("openfga/openfga:v1.14.0") + job := newTestJob(dep) + job.Status = tt.status + dep.Spec.Template.Spec.Containers[0].Env[1].Value = "postgres://db.example.com/openfga" + r := newReconciler(t, nil, dep, job) + + if result := reconcileOnce(t, r); result.RequeueAfter != 10*time.Second { + t.Errorf("expected the job to be polled, got %v", result.RequeueAfter) + } + if _, err := getJob(r); err != nil { + t.Errorf("a started migration must not be replaced: %v", err) + } + }) } } -func TestReconcile_StaleJob_DeletedAndRequeued(t *testing.T) { - // Given: a Deployment at v1.15.0 with an existing migration Job for v1.14.0. - dep := newTestDeployment("openfga", "default", "openfga/openfga:v1.15.0", 0) - dep.Annotations[AnnotationDesiredReplicas] = "3" - - staleJob := &batchv1.Job{ +// The legacy chart's Helm hook Job has the same name and carries the chart's +// app.kubernetes.io/version label, which is the chart appVersion rather than +// the image it ran. It must never be taken as proof of a migration. +func TestReconcile_LegacyHookJob_NotTrusted(t *testing.T) { + legacy := &batchv1.Job{ ObjectMeta: metav1.ObjectMeta{ - Name: "openfga-migrate", - Namespace: "default", + Name: jobKey.Name, + Namespace: jobKey.Namespace, Labels: map[string]string{ - "app.kubernetes.io/version": "v1.14.0", - }, - Annotations: map[string]string{ - "openfga.dev/desired-version": "v1.14.0", - }, - OwnerReferences: []metav1.OwnerReference{ - { - APIVersion: "apps/v1", - Kind: "Deployment", - Name: "openfga", - UID: "test-uid-123", - }, - }, - }, - Spec: batchv1.JobSpec{ - BackoffLimit: ptr.To(int32(3)), - Template: corev1.PodTemplateSpec{ - Spec: corev1.PodSpec{ - Containers: []corev1.Container{{Name: "migrate", Image: "openfga/openfga:v1.14.0"}}, - RestartPolicy: corev1.RestartPolicyNever, - }, - }, - }, - Status: batchv1.JobStatus{ - Succeeded: 1, - Conditions: []batchv1.JobCondition{ - { - Type: batchv1.JobComplete, - Status: corev1.ConditionTrue, - }, + "app.kubernetes.io/version": "v1.14.0", + "app.kubernetes.io/managed-by": "Helm", }, + Annotations: map[string]string{"helm.sh/hook": "post-install, post-upgrade, post-rollback, post-delete"}, }, + Status: batchv1.JobStatus{Conditions: []batchv1.JobCondition{jobCondition(batchv1.JobComplete, time.Now())}}, } + r := newReconciler(t, nil, newTestDeployment("openfga/openfga:v1.14.0"), legacy) - r := newReconciler(dep, staleJob) - - // When: reconciling. - result, err := r.Reconcile(context.Background(), ctrl.Request{ - NamespacedName: types.NamespacedName{Name: "openfga", Namespace: "default"}, - }) - - // Then: no error, requeue to recreate with correct version. - if err != nil { - t.Fatalf("unexpected error: %v", err) + reconcileOnce(t, r) + if _, err := getJob(r); !apierrors.IsNotFound(err) { + t.Errorf("expected the legacy hook job to be deleted, got err=%v", err) } - if result.RequeueAfter == 0 { - t.Error("expected requeue after deleting stale job") + if _, err := getStatus(r); !apierrors.IsNotFound(err) { + t.Errorf("the legacy hook job must not be recorded as a migration, got err=%v", err) } - // Verify the stale Job was deleted. - deletedJob := &batchv1.Job{} - if getErr := r.Get(context.Background(), types.NamespacedName{ - Name: "openfga-migrate", Namespace: "default", - }, deletedJob); getErr == nil { - t.Error("expected stale migration job to be deleted") + reconcileOnce(t, r) + job, err := getJob(r) + if err != nil { + t.Fatalf("expected a new migration job: %v", err) } - - // Verify ConfigMap was NOT updated (migration didn't actually run for v1.15.0). - cm := &corev1.ConfigMap{} - if getErr := r.Get(context.Background(), types.NamespacedName{ - Name: "openfga-migration-status", Namespace: "default", - }, cm); getErr == nil { - if cm.Data["version"] == "v1.15.0" { - t.Error("ConfigMap should not be updated to v1.15.0 from a stale v1.14.0 job") - } + if job.Annotations[AnnotationDesiredVersion] != "v1.14.0" { + t.Errorf("expected the new job to target v1.14.0, got %v", job.Annotations) } } func TestReconcile_MigrationNotEnabled_Skips(t *testing.T) { - // Given: a Deployment without the migration-enabled annotation. - dep := newTestDeployment("openfga", "default", "openfga/openfga:v1.14.0", 3) + dep := newTestDeployment("openfga/openfga:v1.14.0") delete(dep.Annotations, AnnotationMigrationEnabled) + r := newReconciler(t, nil, dep) - r := newReconciler(dep) - - // When: reconciling. - result, err := r.Reconcile(context.Background(), ctrl.Request{ - NamespacedName: types.NamespacedName{Name: "openfga", Namespace: "default"}, - }) - - // Then: no error, no requeue, no Job created, replicas unchanged. - if err != nil { - t.Fatalf("unexpected error: %v", err) + if result := reconcileOnce(t, r); result.RequeueAfter != 0 { + t.Errorf("expected no requeue, got %v", result.RequeueAfter) } - if result.RequeueAfter != 0 { - t.Error("expected no requeue when migration is not enabled") - } - - // Verify no Job was created. - job := &batchv1.Job{} - if getErr := r.Get(context.Background(), types.NamespacedName{ - Name: "openfga-migrate", Namespace: "default", - }, job); getErr == nil { - t.Error("expected no migration job when migration is not enabled") - } - - // Verify replicas unchanged. - updated := &appsv1.Deployment{} - if getErr := r.Get(context.Background(), types.NamespacedName{ - Name: "openfga", Namespace: "default", - }, updated); getErr != nil { - t.Fatalf("getting deployment: %v", getErr) - } - if *updated.Spec.Replicas != 3 { - t.Errorf("expected 3 replicas unchanged, got %d", *updated.Spec.Replicas) + if _, err := getJob(r); !apierrors.IsNotFound(err) { + t.Errorf("expected no migration job, got err=%v", err) } } -func TestReconcile_StaleJob_LabelOnlyFallback_DeletedAndRequeued(t *testing.T) { - // Given: a Deployment at v1.15.0 with an existing Job that only has a label (no annotation). - dep := newTestDeployment("openfga", "default", "openfga/openfga:v1.15.0", 0) - dep.Annotations[AnnotationDesiredReplicas] = "3" - - staleJob := &batchv1.Job{ - ObjectMeta: metav1.ObjectMeta{ - Name: "openfga-migrate", - Namespace: "default", - Labels: map[string]string{ - "app.kubernetes.io/version": "v1.14.0", - }, - // No annotation — forces the label-only fallback path. - OwnerReferences: []metav1.OwnerReference{ - { - APIVersion: "apps/v1", - Kind: "Deployment", - Name: "openfga", - UID: "test-uid-123", - }, - }, - }, - Spec: batchv1.JobSpec{ - BackoffLimit: ptr.To(int32(3)), - Template: corev1.PodTemplateSpec{ - Spec: corev1.PodSpec{ - Containers: []corev1.Container{{Name: "migrate", Image: "openfga/openfga:v1.14.0"}}, - RestartPolicy: corev1.RestartPolicyNever, - }, - }, - }, - } - - r := newReconciler(dep, staleJob) - - // When: reconciling. - result, err := r.Reconcile(context.Background(), ctrl.Request{ - NamespacedName: types.NamespacedName{Name: "openfga", Namespace: "default"}, - }) - - // Then: stale Job should be deleted and requeue requested. - if err != nil { - t.Fatalf("unexpected error: %v", err) - } - if result.RequeueAfter == 0 { - t.Error("expected requeue after deleting stale job") - } - - deletedJob := &batchv1.Job{} - if getErr := r.Get(context.Background(), types.NamespacedName{ - Name: "openfga-migrate", Namespace: "default", - }, deletedJob); getErr == nil { - t.Error("expected stale migration job to be deleted") - } -} - -func TestReconcile_JobSucceeded_UpdatesExistingConfigMap(t *testing.T) { - // Given: a Deployment with a pre-existing ConfigMap from v1.13.0 and a succeeded Job for v1.14.0. - dep := newTestDeployment("openfga", "default", "openfga/openfga:v1.14.0", 0) - dep.Annotations[AnnotationDesiredReplicas] = "3" - - existingCM := &corev1.ConfigMap{ - ObjectMeta: metav1.ObjectMeta{ - Name: "openfga-migration-status", - Namespace: "default", - Labels: map[string]string{ - LabelPartOf: LabelPartOfValue, - LabelComponent: "migration", - "app.kubernetes.io/managed-by": "openfga-operator", - }, - OwnerReferences: []metav1.OwnerReference{ - { - APIVersion: "apps/v1", - Kind: "Deployment", - Name: "openfga", - UID: "test-uid-123", - }, - }, - }, - Data: map[string]string{ - "version": "v1.13.0", - "migratedAt": "2026-04-01T12:00:00Z", - "jobName": "openfga-migrate", - }, - } - - job := &batchv1.Job{ - ObjectMeta: metav1.ObjectMeta{ - Name: "openfga-migrate", - Namespace: "default", - Annotations: map[string]string{ - "openfga.dev/desired-version": "v1.14.0", - }, - OwnerReferences: []metav1.OwnerReference{ - { - APIVersion: "apps/v1", - Kind: "Deployment", - Name: "openfga", - UID: "test-uid-123", - }, - }, - }, - Spec: batchv1.JobSpec{ - BackoffLimit: ptr.To(int32(3)), - Template: corev1.PodTemplateSpec{ - Spec: corev1.PodSpec{ - Containers: []corev1.Container{{Name: "migrate", Image: "openfga/openfga:v1.14.0"}}, - RestartPolicy: corev1.RestartPolicyNever, - }, - }, - }, - Status: batchv1.JobStatus{ - Succeeded: 1, - Conditions: []batchv1.JobCondition{ - { - Type: batchv1.JobComplete, - Status: corev1.ConditionTrue, - }, - }, - }, - } - - r := newReconciler(dep, existingCM, job) - - // When: reconciling. - _, err := r.Reconcile(context.Background(), ctrl.Request{ - NamespacedName: types.NamespacedName{Name: "openfga", Namespace: "default"}, - }) - - // Then: no error. - if err != nil { - t.Fatalf("unexpected error: %v", err) - } - - // Verify ConfigMap was updated to v1.14.0. - cm := &corev1.ConfigMap{} - if getErr := r.Get(context.Background(), types.NamespacedName{ - Name: "openfga-migration-status", Namespace: "default", - }, cm); getErr != nil { - t.Fatalf("expected ConfigMap to exist: %v", getErr) - } - if cm.Data["version"] != "v1.14.0" { - t.Errorf("expected version v1.14.0 in ConfigMap, got %s", cm.Data["version"]) +func TestReconcile_DeploymentNotFound_NoError(t *testing.T) { + r := newReconciler(t, nil) + if result := reconcileOnce(t, r); result.RequeueAfter != 0 { + t.Errorf("expected no requeue, got %v", result.RequeueAfter) } } -func TestReconcile_MigrationNeeded_DoesNotScaleToZero(t *testing.T) { - // Given: a Deployment with replicas > 0 and no migration-status ConfigMap. - // The operator should create the migration Job WITHOUT scaling to zero, - // relying on OpenFGA's built-in schema version check to gate readiness. - dep := newTestDeployment("openfga", "default", "openfga/openfga:v1.14.0", 3) - dep.Annotations = map[string]string{ - AnnotationMigrationEnabled: "true", - AnnotationDesiredReplicas: "3", +func TestReconcile_ContainerFromAnnotation(t *testing.T) { + dep := newTestDeployment("openfga/openfga:v1.14.0") + dep.Annotations[AnnotationContainerName] = "server" + dep.Spec.Template.Spec.Containers = []corev1.Container{ + {Name: "sidecar", Image: "envoyproxy/envoy:v1.30.0"}, + {Name: "server", Image: "openfga/openfga:v1.14.0"}, } + r := newReconciler(t, nil, dep) + reconcileOnce(t, r) - r := newReconciler(dep) - - // When: reconciling. - result, err := r.Reconcile(context.Background(), ctrl.Request{ - NamespacedName: types.NamespacedName{Name: "openfga", Namespace: "default"}, - }) - - // Then: no error, Job created. + job, err := getJob(r) if err != nil { - t.Fatalf("unexpected error: %v", err) - } - if result.RequeueAfter == 0 { - t.Error("expected requeue after creating job") - } - - // Verify Deployment replicas were NOT changed — pods keep running during migration. - updated := &appsv1.Deployment{} - if getErr := r.Get(context.Background(), types.NamespacedName{ - Name: "openfga", Namespace: "default", - }, updated); getErr != nil { - t.Fatalf("getting deployment: %v", getErr) + t.Fatalf("expected migration job to be created: %v", err) } - if *updated.Spec.Replicas != 3 { - t.Errorf("expected replicas to remain at 3, got %d", *updated.Spec.Replicas) + if image := job.Spec.Template.Spec.Containers[0].Image; image != "openfga/openfga:v1.14.0" { + t.Errorf("expected the annotated container's image, got %s", image) } } -func TestReconcile_JobInProgress_Requeues(t *testing.T) { - // Given: a Deployment with a running Job (no conditions set yet). - dep := newTestDeployment("openfga", "default", "openfga/openfga:v1.14.0", 0) - dep.Annotations[AnnotationDesiredReplicas] = "3" +func TestReconcile_ContainerNotFound_ReturnsError(t *testing.T) { + dep := newTestDeployment("openfga/openfga:v1.14.0") + dep.Spec.Template.Spec.Containers[0].Name = "server" + r := newReconciler(t, nil, dep) - job := &batchv1.Job{ - ObjectMeta: metav1.ObjectMeta{ - Name: "openfga-migrate", - Namespace: "default", - Annotations: map[string]string{ - "openfga.dev/desired-version": "v1.14.0", - }, - OwnerReferences: []metav1.OwnerReference{ - { - APIVersion: "apps/v1", - Kind: "Deployment", - Name: "openfga", - UID: "test-uid-123", - }, - }, - }, - Spec: batchv1.JobSpec{ - BackoffLimit: ptr.To(int32(3)), - Template: corev1.PodTemplateSpec{ - Spec: corev1.PodSpec{ - Containers: []corev1.Container{{Name: "migrate", Image: "openfga/openfga:v1.14.0"}}, - RestartPolicy: corev1.RestartPolicyNever, - }, - }, - }, - Status: batchv1.JobStatus{ - Active: 1, - }, - } - - r := newReconciler(dep, job) - - // When: reconciling. - result, err := r.Reconcile(context.Background(), ctrl.Request{ - NamespacedName: types.NamespacedName{Name: "openfga", Namespace: "default"}, - }) - - // Then: no error, requeue after 10s to poll progress. - if err != nil { - t.Fatalf("unexpected error: %v", err) - } - if result.RequeueAfter != 10*time.Second { - t.Errorf("expected 10s requeue for in-progress job, got %v", result.RequeueAfter) - } - - // Verify Deployment still at 0 replicas. - updated := &appsv1.Deployment{} - if getErr := r.Get(context.Background(), types.NamespacedName{ - Name: "openfga", Namespace: "default", - }, updated); getErr != nil { - t.Fatalf("getting deployment: %v", getErr) - } - if *updated.Spec.Replicas != 0 { - t.Errorf("expected 0 replicas while job in progress, got %d", *updated.Spec.Replicas) + if _, err := r.Reconcile(context.Background(), ctrl.Request{NamespacedName: deploymentKey}); err == nil { + t.Fatal("expected an error when the OpenFGA container is missing") } } @@ -1239,15 +515,16 @@ func TestExtractImageTag(t *testing.T) { {"openfga/openfga:v1.14.0", "v1.14.0"}, {"openfga/openfga:latest", "latest"}, {"openfga/openfga", "latest"}, + {"openfga", "latest"}, {"ghcr.io/openfga/openfga:v1.14.0", "v1.14.0"}, {"registry.example.com:5000/openfga/openfga:v1.14.0", "v1.14.0"}, + {"registry.example.com:5000/openfga/openfga", "latest"}, {"openfga/openfga@sha256:abcdef1234567890", "sha256:abcdef1234567890"}, + {"openfga/openfga:v1.14.0@sha256:abcdef1234567890", "sha256:abcdef1234567890"}, } - for _, tt := range tests { t.Run(tt.image, func(t *testing.T) { - got := extractImageTag(tt.image) - if got != tt.expected { + if got := extractImageTag(tt.image); got != tt.expected { t.Errorf("extractImageTag(%q) = %q, want %q", tt.image, got, tt.expected) } }) @@ -1255,7 +532,6 @@ func TestExtractImageTag(t *testing.T) { } func TestClearMigrationFailedConditionIdempotent(t *testing.T) { - // Absent condition: nothing to clear. dep := &appsv1.Deployment{} if clearMigrationFailedCondition(dep) { t.Error("expected no change when the MigrationFailed condition is absent") @@ -1264,11 +540,7 @@ func TestClearMigrationFailedConditionIdempotent(t *testing.T) { t.Errorf("expected no conditions to be added, got %d", len(dep.Status.Conditions)) } - // Condition present and True: clearing flips it to False (a real change). - dep.Status.Conditions = []appsv1.DeploymentCondition{{ - Type: "MigrationFailed", - Status: corev1.ConditionTrue, - }} + dep.Status.Conditions = []appsv1.DeploymentCondition{{Type: "MigrationFailed", Status: corev1.ConditionTrue}} if !clearMigrationFailedCondition(dep) { t.Error("expected a change when clearing a True MigrationFailed condition") } @@ -1278,21 +550,16 @@ func TestClearMigrationFailedConditionIdempotent(t *testing.T) { } transition := cond.LastTransitionTime - // Already False: clearing again must be a no-op and must not advance - // LastTransitionTime — this is what stops the reconcile status write-churn. if clearMigrationFailedCondition(dep) { t.Error("expected no change when the MigrationFailed condition is already False") } - cond = findCondition(dep.Status.Conditions, "MigrationFailed") - if !cond.LastTransitionTime.Equal(&transition) { + if cond := findCondition(dep.Status.Conditions, "MigrationFailed"); !cond.LastTransitionTime.Equal(&transition) { t.Error("LastTransitionTime must not change when the condition is already False") } } func TestSetMigrationFailedConditionIdempotent(t *testing.T) { dep := &appsv1.Deployment{} - - // First set appends the condition. if !setMigrationFailedCondition(dep, "v1.14.0") { t.Error("expected a change when setting MigrationFailed on a fresh deployment") } @@ -1302,46 +569,13 @@ func TestSetMigrationFailedConditionIdempotent(t *testing.T) { } transition := cond.LastTransitionTime - // Re-setting for the same version is a no-op: no LastTransitionTime churn. if setMigrationFailedCondition(dep, "v1.14.0") { t.Error("expected no change when re-setting the same MigrationFailed condition") } - cond = findCondition(dep.Status.Conditions, "MigrationFailed") - if !cond.LastTransitionTime.Equal(&transition) { - t.Error("LastTransitionTime must not change when the condition is unchanged") - } - - // A different version updates the message but does not transition status, - // so LastTransitionTime stays put. if !setMigrationFailedCondition(dep, "v1.15.0") { t.Error("expected a change when the failure message changes") } - cond = findCondition(dep.Status.Conditions, "MigrationFailed") - if !cond.LastTransitionTime.Equal(&transition) { + if cond := findCondition(dep.Status.Conditions, "MigrationFailed"); !cond.LastTransitionTime.Equal(&transition) { t.Error("LastTransitionTime must not change without a status transition") } } - -func TestIsMutableImageReference(t *testing.T) { - tests := []struct { - image string - mutable bool - }{ - {"openfga/openfga:v1.14.0", false}, - {"openfga/openfga:1.14.0", false}, - {"openfga/openfga:v1.14.0-rc1", false}, - {"openfga/openfga@sha256:abcdef1234567890", false}, - {"openfga/openfga:latest", true}, - {"openfga/openfga:v1.14", true}, - {"openfga/openfga", true}, - {"registry.example.com:5000/openfga/openfga:v1.14.0", false}, - {"registry.example.com:5000/openfga/openfga:latest", true}, - } - for _, tt := range tests { - t.Run(tt.image, func(t *testing.T) { - if got := isMutableImageReference(tt.image); got != tt.mutable { - t.Errorf("isMutableImageReference(%q) = %v, want %v", tt.image, got, tt.mutable) - } - }) - } -} diff --git a/operator/tests/README.md b/operator/tests/README.md index 099ec8b6..9a225f7f 100644 --- a/operator/tests/README.md +++ b/operator/tests/README.md @@ -26,7 +26,7 @@ kind load docker-image openfga/openfga-operator:dev ### 1. Happy Path -Deploys OpenFGA with a Postgres instance. The operator should run the migration and scale OpenFGA up within ~30 seconds. +Deploys OpenFGA with a Postgres instance. The operator should run the migration and all OpenFGA pods should become ready within ~30 seconds. ```bash kubectl create namespace openfga-test @@ -89,10 +89,9 @@ helm install openfga-test charts/openfga -n openfga-test \ - Migration Job runs and fails (each pod times out after ~60s) - After 3 failures (backoffLimit), the operator: - Sets `MigrationFailed: True` condition on the Deployment - - Deletes the failed Job - - Creates a fresh Job after a 60-second delay + - Keeps the failed Job for 60 seconds, then replaces it with a fresh one - This cycle repeats indefinitely -- OpenFGA stays at 0/1 throughout — the chart omits `spec.replicas`, so one pod starts (the Kubernetes default) but the readiness gate holds it NotReady, serving no traffic while the migration keeps failing +- OpenFGA stays at 0/3 throughout: the pods cannot reach the database, so they restart and never pass the readiness check **Watch the failure cycle:** @@ -104,8 +103,8 @@ kubectl get deployment openfga-test -n openfga-test \ # Watch operator logs for delete/retry cycle kubectl logs -n openfga-test deployment/openfga-test-openfga-operator -f # Look for: -# "migration job failed, will delete and retry" -# "deleted failed migration job, will retry" +# "migration job failed" +# "retrying migration" # "created migration job" ``` @@ -115,11 +114,11 @@ kubectl logs -n openfga-test deployment/openfga-test-openfga-operator -f kubectl scale deployment openfga-test-postgres -n openfga-test --replicas=1 ``` -**Expected recovery (within ~60s of Postgres becoming ready):** +**Expected recovery:** -- The currently running migration pod connects and succeeds -- Operator updates the ConfigMap with the new version -- Operator scales OpenFGA to 3/3 replicas +- The next migration Job connects and succeeds (within ~60s of Postgres becoming ready) +- Operator updates the ConfigMap with the new version and sets `MigrationFailed: False` +- OpenFGA pods become ready once their restart back-off expires (up to 5 minutes) - `{"status":"SERVING"}` from the health endpoint **Verify recovery:** @@ -159,15 +158,15 @@ helm install openfga-test charts/openfga -n openfga-test \ - Migration Jobs fail repeatedly (DNS resolution fails for `postgres-does-not-exist`) - Operator sets `MigrationFailed: True` on the Deployment -- Operator deletes failed Jobs and retries every ~60 seconds -- OpenFGA stays at 0/1 indefinitely — the single default pod never passes the readiness gate, so it never serves traffic against an unmigrated database +- Operator replaces each failed Job 60 seconds after it fails +- OpenFGA stays at 0/3 indefinitely and never serves traffic This scenario verifies the operator doesn't crash-loop or consume excessive resources when the database is permanently unavailable. **Verify:** ```bash -# OpenFGA at 0/1 (pod NotReady), operator at 1/1 +# OpenFGA at 0/3 (pods NotReady), operator at 1/1 kubectl get deployments -n openfga-test # MigrationFailed condition present From b645b0ec603397b43e289305c0d0769700295ed7 Mon Sep 17 00:00:00 2001 From: SoulPancake <angbpy@gmail.com> Date: Wed, 23 Sep 2026 14:55:04 +0530 Subject: [PATCH 56/70] operator: use the module path the code lives at github.com/openfga/openfga-operator does not exist; the module is at github.com/openfga/helm-charts/operator, so make the path match. --- operator/cmd/main.go | 2 +- operator/go.mod | 2 +- 2 files changed, 2 insertions(+), 2 deletions(-) diff --git a/operator/cmd/main.go b/operator/cmd/main.go index 94c65ce9..783196db 100644 --- a/operator/cmd/main.go +++ b/operator/cmd/main.go @@ -14,7 +14,7 @@ import ( "sigs.k8s.io/controller-runtime/pkg/log/zap" metricsserver "sigs.k8s.io/controller-runtime/pkg/metrics/server" - "github.com/openfga/openfga-operator/internal/controller" + "github.com/openfga/helm-charts/operator/internal/controller" ) var scheme = runtime.NewScheme() diff --git a/operator/go.mod b/operator/go.mod index 72c27f7e..7265283c 100644 --- a/operator/go.mod +++ b/operator/go.mod @@ -1,4 +1,4 @@ -module github.com/openfga/openfga-operator +module github.com/openfga/helm-charts/operator go 1.26.8 From 572749f715eab1c9650de225f9761b13814b2c64 Mon Sep 17 00:00:00 2001 From: SoulPancake <angbpy@gmail.com> Date: Wed, 23 Sep 2026 14:55:04 +0530 Subject: [PATCH 57/70] ci: address actionlint findings in the operator workflows Read the event name and ref from the environment instead of inlining expressions in the script, and drop an unused loop variable. --- .github/workflows/operator.yml | 4 ++-- .github/workflows/test.yml | 4 ++-- 2 files changed, 4 insertions(+), 4 deletions(-) diff --git a/.github/workflows/operator.yml b/.github/workflows/operator.yml index 6a2bc72d..4909f39b 100644 --- a/.github/workflows/operator.yml +++ b/.github/workflows/operator.yml @@ -72,10 +72,10 @@ jobs: - name: Determine push policy id: policy run: | - if [[ "${{ github.event_name }}" == "push" && "${{ github.ref }}" == "refs/heads/main" ]]; then + if [[ "$GITHUB_EVENT_NAME" == "push" && "$GITHUB_REF" == "refs/heads/main" ]]; then echo "push=true" >> "$GITHUB_OUTPUT" echo "mode=main" >> "$GITHUB_OUTPUT" - elif [[ "${{ github.event_name }}" == "workflow_dispatch" && "${{ inputs.push_image }}" == "true" ]]; then + elif [[ "$GITHUB_EVENT_NAME" == "workflow_dispatch" && "${{ inputs.push_image }}" == "true" ]]; then echo "push=true" >> "$GITHUB_OUTPUT" echo "mode=dispatch" >> "$GITHUB_OUTPUT" else diff --git a/.github/workflows/test.yml b/.github/workflows/test.yml index 5efa259a..6810e6da 100644 --- a/.github/workflows/test.yml +++ b/.github/workflows/test.yml @@ -100,7 +100,7 @@ jobs: # Operator must run the migration Job and write ConfigMap at OLD_VER. # Poll because kubectl wait --for=create requires kubectl >=1.31. - for i in $(seq 1 60); do + for _ in $(seq 1 60); do ver=$(kubectl get configmap "${REL}-migration-status" -n "$NS" \ -o jsonpath='{.data.version}' 2>/dev/null || true) if [ "$ver" = "${OLD_VER}" ]; then @@ -124,7 +124,7 @@ jobs: # Operator must detect the version change, delete the stale Job, # run a new migration, and update the ConfigMap to NEW_VER. - for i in $(seq 1 60); do + for _ in $(seq 1 60); do ver=$(kubectl get configmap "${REL}-migration-status" -n "$NS" \ -o jsonpath='{.data.version}' 2>/dev/null || true) if [ "$ver" = "${NEW_VER}" ]; then From 3d544badb91978a633af123304b3c0bad9fbdffb Mon Sep 17 00:00:00 2001 From: SoulPancake <angbpy@gmail.com> Date: Wed, 23 Sep 2026 14:55:05 +0530 Subject: [PATCH 58/70] chart: enable the operator with openfga-operator.enabled, drop migration.enabled The operator was spread over three top-level values: operator.enabled, openfga-operator (subchart values) and migration.enabled. Use openfga-operator.enabled as the dependency condition, which is Helm's convention and how this chart already toggles postgresql and mysql, so the operator's toggle and configuration live under one key. migration.enabled duplicated datastore.applyMigrations: both meant "do not run migrations from this release", and the helper required both. Keep only applyMigrations. migration.serviceAccount stays, since the parent chart creates that service account. --- .github/ci/operator-postgres-values.yaml | 7 +--- charts/openfga-operator/values.schema.json | 1 + charts/openfga-operator/values.yaml | 4 +++ charts/openfga/Chart.lock | 4 +-- charts/openfga/Chart.yaml | 2 +- charts/openfga/ci/operator-mode-values.yaml | 6 +--- charts/openfga/templates/_helpers.tpl | 2 +- charts/openfga/templates/deployment.yaml | 2 +- charts/openfga/templates/job.yaml | 2 +- charts/openfga/templates/rbac.yaml | 2 +- .../openfga/tests/operator_mode_job_test.yaml | 5 ++- .../tests/operator_mode_rbac_test.yaml | 4 +-- .../operator_mode_serviceaccount_test.yaml | 17 ++++----- charts/openfga/tests/operator_mode_test.yaml | 35 +++++++------------ charts/openfga/values.schema.json | 20 +++-------- charts/openfga/values.yaml | 17 ++++----- docs/adr/001-adopt-openfga-operator.md | 18 +++++----- docs/adr/002-operator-managed-migrations.md | 16 ++++----- operator/README.md | 3 +- operator/tests/values-db-outage.yaml | 9 +---- operator/tests/values-happy-path.yaml | 9 +---- operator/tests/values-no-db.yaml | 9 +---- 22 files changed, 70 insertions(+), 124 deletions(-) diff --git a/.github/ci/operator-postgres-values.yaml b/.github/ci/operator-postgres-values.yaml index 842566e0..b4193e20 100644 --- a/.github/ci/operator-postgres-values.yaml +++ b/.github/ci/operator-postgres-values.yaml @@ -3,17 +3,12 @@ # out of charts/openfga/ci/ because chart-testing only installs one version. replicaCount: 1 -operator: - enabled: true - -migration: - enabled: true - datastore: engine: postgres uriSecret: openfga-e2e-postgres-credentials openfga-operator: + enabled: true image: pullPolicy: Never diff --git a/charts/openfga-operator/values.schema.json b/charts/openfga-operator/values.schema.json index 07fbb49c..0315d88b 100644 --- a/charts/openfga-operator/values.schema.json +++ b/charts/openfga-operator/values.schema.json @@ -5,6 +5,7 @@ "global": { "type": "object" }, + "enabled": { "type": "boolean" }, "replicaCount": { "type": "integer", "minimum": 1 diff --git a/charts/openfga-operator/values.yaml b/charts/openfga-operator/values.yaml index be5f05f6..095b6efc 100644 --- a/charts/openfga-operator/values.yaml +++ b/charts/openfga-operator/values.yaml @@ -1,3 +1,7 @@ +# -- Used as the dependency condition by the openfga chart; has no effect when +# this chart is installed on its own. +enabled: true + replicaCount: 1 image: diff --git a/charts/openfga/Chart.lock b/charts/openfga/Chart.lock index 80114538..ac8d7f30 100644 --- a/charts/openfga/Chart.lock +++ b/charts/openfga/Chart.lock @@ -11,5 +11,5 @@ dependencies: - name: openfga-operator repository: file://../openfga-operator version: 0.1.0 -digest: sha256:d502dc105790995a4368a049c0f593820d08f2f82dc9c9a70480a343c7affe8b -generated: "2026-04-10T11:45:16.638975-04:00" +digest: sha256:3df1161dfa820406918b40bb5bada3da6d156a65631f29d72840087d9063db2e +generated: "2026-09-23T14:52:08.781111+05:30" diff --git a/charts/openfga/Chart.yaml b/charts/openfga/Chart.yaml index c68ca136..f299ff30 100644 --- a/charts/openfga/Chart.yaml +++ b/charts/openfga/Chart.yaml @@ -32,4 +32,4 @@ dependencies: - name: openfga-operator version: "0.1.0" repository: "file://../openfga-operator" - condition: operator.enabled + condition: openfga-operator.enabled diff --git a/charts/openfga/ci/operator-mode-values.yaml b/charts/openfga/ci/operator-mode-values.yaml index 02cb06b4..f080f484 100644 --- a/charts/openfga/ci/operator-mode-values.yaml +++ b/charts/openfga/ci/operator-mode-values.yaml @@ -3,15 +3,11 @@ # memory datastore needs no migration, so the Deployment is not opted in to # operator migrations; the Postgres migration path is covered by the operator # E2E step in .github/workflows/test.yml. -operator: - enabled: true - -migration: - enabled: true datastore: engine: memory openfga-operator: + enabled: true image: pullPolicy: Never diff --git a/charts/openfga/templates/_helpers.tpl b/charts/openfga/templates/_helpers.tpl index 90d64bcd..46b8876a 100644 --- a/charts/openfga/templates/_helpers.tpl +++ b/charts/openfga/templates/_helpers.tpl @@ -89,7 +89,7 @@ Create the name of the migration service account to use (operator mode only) Return true if the openfga-operator runs the database migrations for this release */}} {{- define "openfga.operatorMigrations" -}} -{{- if and .Values.operator.enabled .Values.migration.enabled .Values.datastore.applyMigrations (has .Values.datastore.engine (list "postgres" "mysql")) -}} +{{- if and (index .Values "openfga-operator" "enabled") .Values.datastore.applyMigrations (has .Values.datastore.engine (list "postgres" "mysql")) -}} true {{- end -}} {{- end -}} diff --git a/charts/openfga/templates/deployment.yaml b/charts/openfga/templates/deployment.yaml index 80d07da5..6c81a287 100644 --- a/charts/openfga/templates/deployment.yaml +++ b/charts/openfga/templates/deployment.yaml @@ -47,7 +47,7 @@ spec: serviceAccountName: {{ include "openfga.serviceAccountName" . }} securityContext: {{- toYaml .Values.podSecurityContext | nindent 8 }} - {{- $legacyMigrations := and (not .Values.operator.enabled) (has .Values.datastore.engine (list "postgres" "mysql")) }} + {{- $legacyMigrations := and (not (index .Values "openfga-operator" "enabled")) (has .Values.datastore.engine (list "postgres" "mysql")) }} {{ if or (and $legacyMigrations .Values.datastore.applyMigrations .Values.datastore.waitForMigrations) .Values.extraInitContainers }} initContainers: {{- if and $legacyMigrations .Values.datastore.applyMigrations .Values.datastore.waitForMigrations (eq .Values.datastore.migrationType "job") }} diff --git a/charts/openfga/templates/job.yaml b/charts/openfga/templates/job.yaml index d94aef17..05e3d749 100644 --- a/charts/openfga/templates/job.yaml +++ b/charts/openfga/templates/job.yaml @@ -1,4 +1,4 @@ -{{- if and (not .Values.operator.enabled) (has .Values.datastore.engine (list "postgres" "mysql")) .Values.datastore.applyMigrations (eq .Values.datastore.migrationType "job") -}} +{{- if and (not (index .Values "openfga-operator" "enabled")) (has .Values.datastore.engine (list "postgres" "mysql")) .Values.datastore.applyMigrations (eq .Values.datastore.migrationType "job") -}} apiVersion: batch/v1 kind: Job metadata: diff --git a/charts/openfga/templates/rbac.yaml b/charts/openfga/templates/rbac.yaml index 71d3c096..bbb5613d 100644 --- a/charts/openfga/templates/rbac.yaml +++ b/charts/openfga/templates/rbac.yaml @@ -1,4 +1,4 @@ -{{- if and (not .Values.operator.enabled) .Values.serviceAccount.create -}} +{{- if and (not (index .Values "openfga-operator" "enabled")) .Values.serviceAccount.create -}} apiVersion: rbac.authorization.k8s.io/v1 kind: Role metadata: diff --git a/charts/openfga/tests/operator_mode_job_test.yaml b/charts/openfga/tests/operator_mode_job_test.yaml index 31d57607..014e7abe 100644 --- a/charts/openfga/tests/operator_mode_job_test.yaml +++ b/charts/openfga/tests/operator_mode_job_test.yaml @@ -4,8 +4,7 @@ templates: tests: - it: should not render migration job when operator is enabled set: - operator.enabled: true - migration.enabled: true + openfga-operator.enabled: true datastore.engine: postgres datastore.uri: "postgres://localhost/openfga" datastore.applyMigrations: true @@ -16,7 +15,7 @@ tests: - it: should render migration job when operator is disabled set: - operator.enabled: false + openfga-operator.enabled: false datastore.engine: postgres datastore.uri: "postgres://localhost/openfga" datastore.applyMigrations: true diff --git a/charts/openfga/tests/operator_mode_rbac_test.yaml b/charts/openfga/tests/operator_mode_rbac_test.yaml index bb60846c..10ccba0a 100644 --- a/charts/openfga/tests/operator_mode_rbac_test.yaml +++ b/charts/openfga/tests/operator_mode_rbac_test.yaml @@ -4,7 +4,7 @@ templates: tests: - it: should not render legacy RBAC when operator is enabled set: - operator.enabled: true + openfga-operator.enabled: true serviceAccount.create: true asserts: - hasDocuments: @@ -12,7 +12,7 @@ tests: - it: should render legacy RBAC when operator is disabled set: - operator.enabled: false + openfga-operator.enabled: false serviceAccount.create: true asserts: - hasDocuments: diff --git a/charts/openfga/tests/operator_mode_serviceaccount_test.yaml b/charts/openfga/tests/operator_mode_serviceaccount_test.yaml index 304f4ae7..4d50202e 100644 --- a/charts/openfga/tests/operator_mode_serviceaccount_test.yaml +++ b/charts/openfga/tests/operator_mode_serviceaccount_test.yaml @@ -4,8 +4,7 @@ templates: tests: - it: should render migration service account when operator is enabled set: - operator.enabled: true - migration.enabled: true + openfga-operator.enabled: true datastore.engine: postgres migration.serviceAccount.create: true serviceAccount.create: true @@ -22,7 +21,7 @@ tests: - it: should not render migration service account when operator is disabled set: - operator.enabled: false + openfga-operator.enabled: false serviceAccount.create: true asserts: - hasDocuments: @@ -30,8 +29,7 @@ tests: - it: should not render migration service account when migration SA creation is disabled set: - operator.enabled: true - migration.enabled: true + openfga-operator.enabled: true datastore.engine: postgres migration.serviceAccount.create: false migration.serviceAccount.name: external-sa @@ -42,8 +40,7 @@ tests: - it: should render migration service account with custom annotations set: - operator.enabled: true - migration.enabled: true + openfga-operator.enabled: true datastore.engine: postgres migration.serviceAccount.create: true migration.serviceAccount.annotations: @@ -57,8 +54,7 @@ tests: - it: should use custom migration service account name set: - operator.enabled: true - migration.enabled: true + openfga-operator.enabled: true datastore.engine: postgres migration.serviceAccount.create: true migration.serviceAccount.name: my-migrator @@ -71,8 +67,7 @@ tests: - it: should not render migration service account for the memory datastore set: - operator.enabled: true - migration.enabled: true + openfga-operator.enabled: true migration.serviceAccount.create: true serviceAccount.create: true asserts: diff --git a/charts/openfga/tests/operator_mode_test.yaml b/charts/openfga/tests/operator_mode_test.yaml index cefcbd08..16dbc197 100644 --- a/charts/openfga/tests/operator_mode_test.yaml +++ b/charts/openfga/tests/operator_mode_test.yaml @@ -5,8 +5,7 @@ tests: # --- Deployment annotations --- - it: should set operator annotations when operator and migration are enabled set: - operator.enabled: true - migration.enabled: true + openfga-operator.enabled: true datastore.engine: postgres asserts: - equal: @@ -21,7 +20,7 @@ tests: - it: should not set operator annotations when operator is disabled set: - operator.enabled: false + openfga-operator.enabled: false annotations: custom: value asserts: @@ -33,8 +32,7 @@ tests: - it: should not set operator annotations for the memory datastore set: - operator.enabled: true - migration.enabled: true + openfga-operator.enabled: true datastore.engine: memory asserts: - isNull: @@ -42,8 +40,7 @@ tests: - it: should use custom migration service account name when set set: - operator.enabled: true - migration.enabled: true + openfga-operator.enabled: true datastore.engine: postgres migration.serviceAccount.name: my-custom-sa asserts: @@ -53,8 +50,7 @@ tests: - it: should not set migration-service-account annotation when SA creation is disabled and no name set set: - operator.enabled: true - migration.enabled: true + openfga-operator.enabled: true datastore.engine: postgres migration.serviceAccount.create: false asserts: @@ -65,8 +61,7 @@ tests: # The operator never changes the replica count, so it renders as in legacy mode. - it: should set replicas to replicaCount in operator mode set: - operator.enabled: true - migration.enabled: true + openfga-operator.enabled: true replicaCount: 3 datastore.engine: postgres asserts: @@ -76,8 +71,7 @@ tests: - it: should leave replicas to the autoscaler in operator mode set: - operator.enabled: true - migration.enabled: true + openfga-operator.enabled: true autoscaling.enabled: true datastore.engine: postgres asserts: @@ -89,7 +83,7 @@ tests: - it: should set replicas to replicaCount when operator is disabled set: - operator.enabled: false + openfga-operator.enabled: false replicaCount: 5 datastore.engine: postgres asserts: @@ -100,8 +94,7 @@ tests: # applyMigrations=false opts out: no operator annotations, replicas rendered normally. - it: should not use operator mode when applyMigrations is false set: - operator.enabled: true - migration.enabled: true + openfga-operator.enabled: true replicaCount: 4 datastore.engine: postgres datastore.applyMigrations: false @@ -115,8 +108,7 @@ tests: # --- initContainers gating --- - it: should not render migration initContainers when operator is enabled set: - operator.enabled: true - migration.enabled: true + openfga-operator.enabled: true datastore.engine: postgres datastore.uri: "postgres://localhost/openfga" datastore.applyMigrations: true @@ -128,7 +120,7 @@ tests: - it: should render migration initContainers when operator is disabled set: - operator.enabled: false + openfga-operator.enabled: false datastore.engine: postgres datastore.uri: "postgres://localhost/openfga" datastore.applyMigrations: true @@ -145,7 +137,7 @@ tests: # across upgrades. Regression guard for the operator-migration branch. - it: should include common labels on pod template metadata when operator is disabled set: - operator.enabled: false + openfga-operator.enabled: false asserts: - isNotEmpty: path: spec.template.metadata.labels["helm.sh/chart"] @@ -158,8 +150,7 @@ tests: - it: should include common labels on pod template metadata when operator is enabled set: - operator.enabled: true - migration.enabled: true + openfga-operator.enabled: true datastore.engine: postgres asserts: - isNotEmpty: diff --git a/charts/openfga/values.schema.json b/charts/openfga/values.schema.json index 07d2d2ec..bc56e79a 100644 --- a/charts/openfga/values.schema.json +++ b/charts/openfga/values.schema.json @@ -1293,31 +1293,21 @@ "description": "This value is not used by this chart, but allows a common pattern of enabling/disabling subchart dependencies (where OpenFGA is a subchart)", "default": false }, - "operator": { + "openfga-operator": { "type": "object", - "description": "Controls the openfga-operator subchart. When enabled, migration is managed by the operator instead of the Helm job hook.", + "description": "Configuration for the openfga-operator subchart, validated by that chart's own schema. When enabled, database migrations are run by the operator instead of the Helm hook Job.", "properties": { "enabled": { "type": "boolean", - "description": "Enable the openfga-operator subchart for operator-managed migrations", + "description": "Install the openfga-operator and let it run database migrations", "default": false } - }, - "additionalProperties": false - }, - "openfga-operator": { - "type": "object", - "description": "Values passed through to the openfga-operator subchart (validated by that chart's own schema)" + } }, "migration": { "type": "object", - "description": "Controls operator-driven migration behavior. Only used when operator.enabled is true.", + "description": "Service account for the migration Jobs the operator creates. Only used when openfga-operator.enabled is true.", "properties": { - "enabled": { - "type": "boolean", - "description": "Enable operator-managed database migrations", - "default": true - }, "serviceAccount": { "type": "object", "properties": { diff --git a/charts/openfga/values.yaml b/charts/openfga/values.yaml index 94be1ded..acef64f1 100644 --- a/charts/openfga/values.yaml +++ b/charts/openfga/values.yaml @@ -386,14 +386,11 @@ testContainerSpec: {} ## Note: Supports use of custom Helm templates extraObjects: [] -# -- operator controls the openfga-operator subchart. -# When enabled, migration is managed by the operator instead of the Helm job hook. -operator: - enabled: false - -# -- Values passed to the openfga-operator subchart (when operator.enabled is true). +# -- Configuration for the openfga-operator subchart. When enabled, database +# migrations are run by the operator instead of the Helm hook Job. # See charts/openfga-operator/values.yaml for all available options. -openfga-operator: {} +openfga-operator: + enabled: false # migrationJob: # backoffLimit: 3 # activeDeadlineSeconds: 0 @@ -408,11 +405,9 @@ openfga-operator: {} # limits: # memory: 128Mi -# -- migration controls operator-driven migration behavior. -# Only used when operator.enabled is true. +# -- Service account for the migration Jobs the operator creates. +# Only used when openfga-operator.enabled is true. migration: - # -- Enable operator-managed migrations. Set to false if you manage migrations externally. - enabled: true serviceAccount: # -- Create a dedicated service account for migration Jobs. # The migration Job inherits env vars (including secretKeyRef) from the OpenFGA container. diff --git a/docs/adr/001-adopt-openfga-operator.md b/docs/adr/001-adopt-openfga-operator.md index 63ee980f..35c58ddf 100644 --- a/docs/adr/001-adopt-openfga-operator.md +++ b/docs/adr/001-adopt-openfga-operator.md @@ -60,7 +60,7 @@ We will build an **OpenFGA Kubernetes Operator** that handles: The operator will be: - Written in Go using `controller-runtime` / kubebuilder - Distributed as a Helm subchart dependency of the main OpenFGA chart -- Optional — users who don't need it can set `operator.enabled: false` and fall back to the existing behavior +- Optional — users who don't need it can set `openfga-operator.enabled: false` and fall back to the existing behavior Development will follow a staged approach to deliver value incrementally: @@ -77,30 +77,30 @@ Stage 1 has shipped on the `feat/operator-migration` branch. Stages 2-4 are plan ### Delivered in Stage 1 -- Operator Go project under `/operator/`, built with `controller-runtime` and kubebuilder scaffolding -- Operator packaged as a Helm subchart (`charts/openfga-operator/`) and wired into the main chart via a `condition: operator.enabled` dependency -- `operator.enabled` values toggle (default `false`) that gates all operator-managed behavior +- Operator Go project under `/operator/`, built with `controller-runtime` +- Operator packaged as a Helm subchart (`charts/openfga-operator/`) and wired into the main chart via a `condition: openfga-operator.enabled` dependency +- `openfga-operator.enabled` values toggle (default `false`) that gates all operator-managed behavior - Migration reconciler (`migration_controller.go`) that runs migration Jobs when the operator is enabled - Separate migration ServiceAccount with IAM-annotation support (`openfga.migrationServiceAccountName` helper), created when the operator is enabled ### Deferred to later stages -- `FGAStore`, `FGAModel`, and `FGATuples` CRDs and their controllers — `charts/openfga-operator/crds/` is reserved but intentionally empty in Stage 1 +- `FGAStore`, `FGAModel`, and `FGATuples` CRDs and their controllers - Declarative store/model/tuple lifecycle management ### Backward-compatibility path (deprecated) -When `operator.enabled: false`, the chart still renders the legacy migration path: the Helm-hook migration Job, the `groundnuty/k8s-wait-for` init container, and the job-status RBAC. **This path is deprecated and will be removed in a future release** once the operator is the default and users have had time to migrate. It remains only to preserve backward compatibility during the transition. +When `openfga-operator.enabled: false`, the chart still renders the legacy migration path: the Helm-hook migration Job, the `groundnuty/k8s-wait-for` init container, and the job-status RBAC. **This path is deprecated and will be removed in a future release** once the operator is the default and users have had time to migrate. It remains only to preserve backward compatibility during the transition. ## Consequences ### Positive - **Resolves all 6 migration issues** (#211, #107, #120, #100, #95, #126) and related dependency issues (#132, #144) on the operator-enabled path -- **Removes `k8s-wait-for` from the operator-enabled path** — the unmaintained, CVE-carrying image is no longer used when `operator.enabled: true`, and will be removed from the chart entirely once the legacy path is retired +- **Removes `k8s-wait-for` from the operator-enabled path** — the unmaintained, CVE-carrying image is no longer used when `openfga-operator.enabled: true`, and will be removed from the chart entirely once the legacy path is retired - **Enables GitOps-native authorization management** (planned, Stages 2-4) — stores, models, and tuples will become declarative Kubernetes resources that ArgoCD/FluxCD can sync - **Enforces least-privilege** — separate ServiceAccounts for migration (DDL) and runtime (CRUD) on the operator-enabled path -- **Path to simplifying the Helm chart** — the migration Job template, init container logic, job-status RBAC, and hook annotations are conditionalized behind `operator.enabled: false` and scheduled for removal when the legacy path is retired +- **Path to simplifying the Helm chart** — the migration Job template, init container logic, job-status RBAC, and hook annotations are conditionalized behind `openfga-operator.enabled: false` and scheduled for removal when the legacy path is retired - **Follows Kubernetes ecosystem conventions** — operators are the standard pattern for managing stateful application lifecycle ### Negative @@ -113,5 +113,5 @@ When `operator.enabled: false`, the chart still renders the legacy migration pat ### Neutral -- **Backward compatibility preserved during the transition** — `operator.enabled: false` keeps the existing Helm-hook behavior working for users who have not yet migrated, but this path is deprecated and slated for removal +- **Backward compatibility preserved during the transition** — `openfga-operator.enabled: false` keeps the existing Helm-hook behavior working for users who have not yet migrated, but this path is deprecated and slated for removal - **No change for memory-datastore users** — users running with `datastore.engine: memory` are unaffected (no migrations, no operator needed) diff --git a/docs/adr/002-operator-managed-migrations.md b/docs/adr/002-operator-managed-migrations.md index 5f35a0cb..e51923be 100644 --- a/docs/adr/002-operator-managed-migrations.md +++ b/docs/adr/002-operator-managed-migrations.md @@ -186,9 +186,9 @@ No hooks. No init containers. No `k8s-wait-for`. All resources are regular Kuber ### What Changes in the Helm Chart -Nothing is deleted outright — every change is gated on `operator.enabled` so the legacy flow remains the default for backward compatibility. +Nothing is deleted outright — every change is gated on `openfga-operator.enabled` so the legacy flow remains the default for backward compatibility. -**Gated on `operator.enabled: false` (legacy Helm-hook flow, rendered when the operator is disabled):** +**Gated on `openfga-operator.enabled: false` (legacy Helm-hook flow, rendered when the operator is disabled):** | File/Section | Behavior when operator is enabled | |--------------|-----------------------------------| @@ -199,17 +199,17 @@ Nothing is deleted outright — every change is gated on `operator.enabled` so t | `values.yaml`: `migrate.annotations` | Unused — no Helm hooks | | Deployment migration init containers | Skipped — OpenFGA's readiness check holds pods until the schema is migrated | -**Added (active only when `operator.enabled: true`):** +**Added (active only when `openfga-operator.enabled: true`):** | File/Section | Purpose | |--------------|---------| -| `values.yaml`: `operator.enabled` | Toggle the operator subchart | -| `values.yaml`: `migration.serviceAccount.*` | Separate ServiceAccount for migration Jobs | +| `values.yaml`: `openfga-operator.enabled` | Toggle the operator subchart | | `values.yaml`: `openfga-operator.migrationJob.*` | Migration Job backoff, deadline, and TTL configuration | +| `values.yaml`: `migration.serviceAccount.*` | Separate ServiceAccount for migration Jobs | | `templates/serviceaccount.yaml`: second SA | Migration ServiceAccount | | `charts/openfga-operator/` | Operator subchart (conditional dependency) | -Users on `operator.enabled: false` (the default) see identical rendered output to the pre-operator chart, so gradual adoption is possible with no forced migration. +Users on `openfga-operator.enabled: false` (the default) see identical rendered output to the pre-operator chart, so gradual adoption is possible with no forced migration. ## Consequences @@ -218,14 +218,14 @@ Users on `operator.enabled: false` (the default) see identical rendered output t - **All 6 migration issues resolved** — no Helm hooks means no ArgoCD/FluxCD/`--wait` incompatibility - **`k8s-wait-for` eliminated** — removes an unmaintained image with CVEs from the supply chain (#132, #144) - **Least-privilege enforced** — separate ServiceAccounts for migration (DDL) and runtime (CRUD) (#95) -- **Runtime surface area reduced** — when `operator.enabled: true`, the legacy migration Job, init-container `k8s-wait-for` logic, and job-watching RBAC are skipped from the rendered manifest +- **Runtime surface area reduced** — when `openfga-operator.enabled: true`, the legacy migration Job, init-container `k8s-wait-for` logic, and job-watching RBAC are skipped from the rendered manifest - **Migration is observable** — Job is a regular resource visible in all tools; ConfigMap records migration history; operator conditions surface errors - **Idempotent and crash-safe** — operator can restart at any point and resume correctly ### Negative - **Operator is a new runtime dependency** — if the operator pod is unavailable, migrations don't run (but existing running pods are unaffected) -- **Two upgrade paths to document** — `operator.enabled: true` (new) vs `operator.enabled: false` (legacy) +- **Two upgrade paths to document** — `openfga-operator.enabled: true` (new) vs `openfga-operator.enabled: false` (legacy) ### Risks diff --git a/operator/README.md b/operator/README.md index 6a306d93..fea87c62 100644 --- a/operator/README.md +++ b/operator/README.md @@ -124,7 +124,7 @@ The operator reads these annotations from the OpenFGA Deployment: | Annotation | Description | |------------|-------------| -| `openfga.dev/migration-enabled` | Must be `"true"` for the operator to manage migrations. Deployments without this annotation are ignored. Set by the Helm chart when `operator.enabled`, `migration.enabled`, and `datastore.applyMigrations` are true and the datastore is Postgres or MySQL. | +| `openfga.dev/migration-enabled` | Must be `"true"` for the operator to manage migrations. Deployments without this annotation are ignored. Set by the Helm chart when `openfga-operator.enabled` and `datastore.applyMigrations` are true and the datastore is Postgres or MySQL. | | `openfga.dev/container-name` | The OpenFGA container in the pod spec. Defaults to `openfga`. | | `openfga.dev/migration-service-account` | The ServiceAccount to use for migration Jobs. Defaults to the Deployment's SA. | @@ -133,4 +133,5 @@ The operator reads these annotations from the OpenFGA Deployment: - **Migrations key only on the image tag:** The operator compares the container image tag (or digest) to the `{name}-migration-status` ConfigMap. A mutable tag like `latest`, or a tag reused for a new build, is not seen as a change, so the migration is skipped — use immutable tags (e.g. `v1.14.0`) or pin by digest. A migration-needing change that keeps the same image — for example repointing `datastore.uri` at a different or restored database — also won't trigger a Job; delete the status ConfigMap (and the `{name}-migrate` Job, if it still exists) to run the migration again. - **Legacy migration values:** `migrate.*` (extra volumes and mounts, init containers, sidecars, annotations, labels, timeout) and `datastore.migrations.resources` only apply to the legacy Helm hook Job. The operator's Job copies the OpenFGA container's image, env, volumes, resources, security context and scheduling instead, so put anything the migration needs (e.g. CA bundles) in the top-level `extraVolumes`, `extraVolumeMounts` and `extraEnvVars`. - **Single-container migration Job:** The Job runs one container (`openfga migrate`) and injects no sidecars or extra init containers, so databases reached through a sidecar proxy (Cloud SQL Auth Proxy, AlloyDB) aren't supported for operator-managed migrations. A sidecar injected into every pod in the namespace that doesn't exit on its own (e.g. an Istio sidecar) keeps the Job pod running and stops the Job from completing. +- **Job pod labels:** The migration pod is labelled `app.kubernetes.io/part-of: openfga` and `app.kubernetes.io/component: migration`, not with the OpenFGA Deployment's `app.kubernetes.io/name`/`instance` labels (which would make it a Service endpoint). A NetworkPolicy that allows database egress only for the OpenFGA pods' labels needs a rule for the migration pod too. - **One namespace per operator:** The operator reconciles every opted-in OpenFGA Deployment in its watch namespace. Operators installed by several releases in one namespace share a leader election lease, so only one of them is active at a time. diff --git a/operator/tests/values-db-outage.yaml b/operator/tests/values-db-outage.yaml index a7c59720..16d7427e 100644 --- a/operator/tests/values-db-outage.yaml +++ b/operator/tests/values-db-outage.yaml @@ -1,8 +1,6 @@ # Test values: Postgres deployed but scaled to 0 (simulates DB outage) -operator: - enabled: true - openfga-operator: + enabled: true image: repository: openfga/openfga-operator tag: dev @@ -16,11 +14,6 @@ datastore: engine: postgres uri: "postgres://openfga:changeme@openfga-test-postgres:5432/openfga?sslmode=disable" -migration: - enabled: true - serviceAccount: - create: true - extraObjects: - apiVersion: v1 kind: Secret diff --git a/operator/tests/values-happy-path.yaml b/operator/tests/values-happy-path.yaml index 77e6306f..53f32717 100644 --- a/operator/tests/values-happy-path.yaml +++ b/operator/tests/values-happy-path.yaml @@ -1,8 +1,6 @@ # Local test values for operator-managed migration on Rancher Desktop -operator: - enabled: true - openfga-operator: + enabled: true image: repository: openfga/openfga-operator tag: dev @@ -16,11 +14,6 @@ datastore: engine: postgres uri: "postgres://openfga:changeme@openfga-test-postgres:5432/openfga?sslmode=disable" -migration: - enabled: true - serviceAccount: - create: true - extraObjects: - apiVersion: v1 kind: Secret diff --git a/operator/tests/values-no-db.yaml b/operator/tests/values-no-db.yaml index 2d1cd762..e43000cd 100644 --- a/operator/tests/values-no-db.yaml +++ b/operator/tests/values-no-db.yaml @@ -1,8 +1,6 @@ # Test values with NO postgres — simulates database unavailable -operator: - enabled: true - openfga-operator: + enabled: true image: repository: openfga/openfga-operator tag: dev @@ -16,8 +14,3 @@ datastore: engine: postgres # Points to a service that doesn't exist uri: "postgres://openfga:changeme@postgres-does-not-exist:5432/openfga?sslmode=disable" - -migration: - enabled: true - serviceAccount: - create: true From bccdb31f2f9cd3467d2136e98096a61166f9e17f Mon Sep 17 00:00:00 2001 From: SoulPancake <angbpy@gmail.com> Date: Wed, 23 Sep 2026 14:55:05 +0530 Subject: [PATCH 59/70] docs: document operator mode in the chart README and tidy the ADRs Add a section on operator-run migrations to the openfga chart README and list the openfga-operator chart in the repository README, since chart-releaser publishes it. Note that the migration pod does not carry the OpenFGA pod labels, which matters for NetworkPolicies. Make the ADR index match ADR-001's status, drop the claim that the operator was scaffolded with kubebuilder, and remove the empty crds/ placeholder directory from the operator chart. --- README.md | 1 + charts/openfga-operator/crds/README.md | 4 ---- charts/openfga/README.md | 15 +++++++++++++++ docs/adr/README.md | 2 +- 4 files changed, 17 insertions(+), 5 deletions(-) delete mode 100644 charts/openfga-operator/crds/README.md diff --git a/README.md b/README.md index ad9446ee..5721bce4 100644 --- a/README.md +++ b/README.md @@ -14,6 +14,7 @@ It is designed to make it easy for developers to model their application permiss ## Charts * [openfga](https://github.com/openfga/helm-charts/blob/main/charts/openfga) +* [openfga-operator](https://github.com/openfga/helm-charts/blob/main/charts/openfga-operator) — runs OpenFGA database migrations; installed by the openfga chart with `openfga-operator.enabled: true` ## Contributing diff --git a/charts/openfga-operator/crds/README.md b/charts/openfga-operator/crds/README.md deleted file mode 100644 index 060b0d0c..00000000 --- a/charts/openfga-operator/crds/README.md +++ /dev/null @@ -1,4 +0,0 @@ -# CRDs - -This directory is reserved for Custom Resource Definitions added in later stages. -No CRDs are installed in Stage 1 (migration orchestration). diff --git a/charts/openfga/README.md b/charts/openfga/README.md index a85e3815..f9ae8971 100644 --- a/charts/openfga/README.md +++ b/charts/openfga/README.md @@ -151,6 +151,21 @@ datastore: passwordKey: password ``` +### Running migrations with the operator + +By default the chart runs database migrations from a Helm hook Job and gates the OpenFGA pods on it with an init container. Helm hooks are not run by Argo CD and conflict with `helm install --wait` and Flux, so the chart can instead install the [openfga-operator](../openfga-operator), which runs `openfga migrate` as a regular Job whenever the OpenFGA image changes: + +```yaml +openfga-operator: + enabled: true + +datastore: + engine: postgres + uriSecret: my-postgres-secret +``` + +The operator only runs migrations; replicas, autoscaling and the pod template stay under the chart's control. It records the migrated version in the `<release>-migration-status` ConfigMap and sets a `MigrationFailed` condition on the Deployment if a migration fails. The migration Job runs as a dedicated `<release>-migration` service account (`migration.serviceAccount`), which can carry cloud IAM annotations for DDL permissions. Migrations only run when the image tag changes, so pin `image.tag` to a release rather than a floating tag. See the [operator README](../../operator/README.md) for how it works and its limitations. + ## Uninstalling the Chart To uninstall/delete the `openfga` deployment: diff --git a/docs/adr/README.md b/docs/adr/README.md index d6b3445e..536a9d02 100644 --- a/docs/adr/README.md +++ b/docs/adr/README.md @@ -10,7 +10,7 @@ We follow the format described by [Michael Nygard](https://cognitect.com/blog/20 | ADR | Title | Status | Date | |-----|-------|--------|------| -| [ADR-001](001-adopt-openfga-operator.md) | Adopt a Kubernetes Operator for OpenFGA Lifecycle Management | Proposed | 2026-04-06 | +| [ADR-001](001-adopt-openfga-operator.md) | Adopt a Kubernetes Operator for OpenFGA Lifecycle Management | Accepted | 2026-04-06 | | [ADR-002](002-operator-managed-migrations.md) | Replace Helm Hook Migrations with Operator-Managed Migrations | Proposed | 2026-04-06 | --- From 1655b0c48388549b04ccf13d09a63021815b09a2 Mon Sep 17 00:00:00 2001 From: SoulPancake <angbpy@gmail.com> Date: Wed, 23 Sep 2026 15:18:35 +0530 Subject: [PATCH 60/70] ci: test operator changes against a cluster and guard the image version - Run the kind, ct install and E2E steps when operator/ changes, not only when a chart changes; operator code was otherwise never tested against a cluster on its own. - Fail operator PRs that change the image inputs without bumping the operator chart's appVersion and version, since CI publishes each version tag once. Document the full release chain in the operator README. - Upgrade to the chart's appVersion in the E2E instead of a pinned v1.14.1, so every OpenFGA version bump exercises migrating to it. - Run the operator's unit tests with -race. --- .github/workflows/operator.yml | 30 +++++++++++++++++++++++++++++- .github/workflows/test.yml | 12 ++++++++---- operator/README.md | 7 +++++++ 3 files changed, 44 insertions(+), 5 deletions(-) diff --git a/.github/workflows/operator.yml b/.github/workflows/operator.yml index 4909f39b..bc1f79d9 100644 --- a/.github/workflows/operator.yml +++ b/.github/workflows/operator.yml @@ -31,6 +31,34 @@ jobs: steps: - name: Checkout uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1 + with: + fetch-depth: 0 + + - name: Require a version bump when the operator image changes + if: github.event_name == 'pull_request' + env: + BASE: ${{ github.base_ref }} + run: | + set -euo pipefail + # CI publishes ghcr.io/openfga/openfga-operator:<appVersion> once and never + # overwrites it, so a change to the image is only released when the + # operator chart's appVersion and version are bumped. + image_inputs=(operator/cmd operator/internal operator/go.mod operator/go.sum operator/Dockerfile) + if git diff --quiet "origin/$BASE...HEAD" -- "${image_inputs[@]}"; then + echo "operator image inputs unchanged" + exit 0 + fi + field() { grep "^$1:" "$2" | awk '{print $2}' | tr -d '"'; } + if ! git show "origin/$BASE:charts/openfga-operator/Chart.yaml" > /tmp/base-chart.yaml 2>/dev/null; then + echo "operator chart is new in this PR" + exit 0 + fi + for f in appVersion version; do + if [[ "$(field $f /tmp/base-chart.yaml)" == "$(field $f charts/openfga-operator/Chart.yaml)" ]]; then + echo "::error file=charts/openfga-operator/Chart.yaml::operator image inputs changed but $f is still $(field $f charts/openfga-operator/Chart.yaml); bump appVersion and version so the change is published" + exit 1 + fi + done - name: Set up Go uses: actions/setup-go@40f1582b2485089dde7abd97c1529aa768e1baff # v5.6.0 @@ -40,7 +68,7 @@ jobs: - name: Run tests working-directory: operator - run: go test ./... -v + run: go test ./... -race -v - name: Check formatting working-directory: operator diff --git a/.github/workflows/test.yml b/.github/workflows/test.yml index 6810e6da..04d82919 100644 --- a/.github/workflows/test.yml +++ b/.github/workflows/test.yml @@ -38,7 +38,10 @@ jobs: id: list-changed run: | changed=$(ct list-changed --target-branch ${{ github.event.repository.default_branch }}) - if [[ -n "$changed" ]]; then + # Operator code changes are only exercised against a cluster by the + # steps below, so treat them as a chart change too. + operator=$(git diff --name-only "origin/${{ github.event.repository.default_branch }}...HEAD" -- operator) + if [[ -n "$changed" || -n "$operator" ]]; then echo "changed=true" >> "$GITHUB_OUTPUT" fi @@ -77,12 +80,13 @@ jobs: env: NS: openfga-e2e REL: openfga - # v1.9.5 → v1.14.1 crosses the v1.10.0 "!!REQUIRES MIGRATION!!" - # boundary (collation spec change in openfga/openfga#2661). + # v1.9.5 predates the v1.10.0 "!!REQUIRES MIGRATION!!" boundary + # (collation spec change in openfga/openfga#2661), so upgrading to the + # chart's appVersion always crosses at least one migration. OLD_VER: v1.9.5 - NEW_VER: v1.14.1 run: | set -euo pipefail + NEW_VER=$(grep '^appVersion:' charts/openfga/Chart.yaml | awk '{print $2}' | tr -d '"') kubectl create namespace "$NS" helm dependency build charts/openfga diff --git a/operator/README.md b/operator/README.md index fea87c62..9da24f71 100644 --- a/operator/README.md +++ b/operator/README.md @@ -49,6 +49,13 @@ go vet ./... docker build -t openfga/openfga-operator:dev . ``` +## Releasing + +CI publishes `ghcr.io/openfga/openfga-operator:<appVersion>` on the first push to `main` that carries that appVersion and never overwrites it, and chart-releaser likewise skips chart versions that already exist. A change to the operator image (`cmd/`, `internal/`, `go.mod`, `go.sum`, `Dockerfile`) therefore has to bump, in the same PR: + +1. `appVersion` and `version` in `charts/openfga-operator/Chart.yaml` (the operator workflow fails the PR otherwise) +2. the `openfga-operator` dependency version and `version` in `charts/openfga/Chart.yaml`, then `helm dependency update charts/openfga` to refresh `Chart.lock` (`helm dependency build` fails otherwise) + ## Local Testing Integration test values and instructions are in [`tests/`](tests/). Three scenarios are provided: From fe2a3c73c9c6e39a0cf9b9b50e8f96945cb148ab Mon Sep 17 00:00:00 2001 From: SoulPancake <angbpy@gmail.com> Date: Wed, 23 Sep 2026 16:15:55 +0530 Subject: [PATCH 61/70] operator: never interrupt a running migration when the image changes A version change replaced the migration Job at once, even with its pod running. Aborting a non-transactional step such as the concurrent index build in Postgres migration 006 leaves an invalid index, and the rerun's IF NOT EXISTS then skips it while goose records the version as applied. Wait for a running Job to finish and replace it afterwards, as already done for pod template changes. --- docs/adr/002-operator-managed-migrations.md | 1 + operator/README.md | 2 + .../controller/migration_controller.go | 15 +++++-- .../controller/migration_controller_test.go | 44 +++++++++++++++++++ 4 files changed, 58 insertions(+), 4 deletions(-) diff --git a/docs/adr/002-operator-managed-migrations.md b/docs/adr/002-operator-managed-migrations.md index e51923be..72e94bc4 100644 --- a/docs/adr/002-operator-managed-migrations.md +++ b/docs/adr/002-operator-managed-migrations.md @@ -122,6 +122,7 @@ The Job created by the operator has no Helm hook annotations. It is a standard K |---------|----------| | Job fails | Operator sets `MigrationFailed` on the Deployment, keeps the failed Job for 60 seconds so its logs can be read, then replaces it. On a fresh database the pods stay `NotReady`; on an upgrade they keep serving on the previous schema. | | Job pod never starts | A bad secret reference, image pull error or unschedulable pod never fails the Job. Once the Deployment's pod template changes (the fix rolls out), the operator rebuilds a Job whose pod is not running. | +| Image changes while a Job runs | The running Job is left to finish and then replaced by one for the new image. The hook flow deletes the running hook Job instead (`before-hook-creation`), which can abort a concurrent index build and leave it invalid. | | Job hangs | No deadline by default, like the Helm hook Job. `activeDeadlineSeconds` can be set, but a migration cut off halfway (an index build, a MySQL table rebuild) starts over on the next attempt. | | Operator crashes | On restart, re-reads the ConfigMap and Job status and resumes. The retry delay is measured from the failed Job's condition, so it survives restarts. | | Database unreachable | Job fails to connect. After exhausting `backoffLimit` the cycle above repeats until the database becomes available. | diff --git a/operator/README.md b/operator/README.md index 9da24f71..24778f3d 100644 --- a/operator/README.md +++ b/operator/README.md @@ -13,6 +13,8 @@ This is **Stage 1** of the operator — focused solely on migration orchestratio - Updates the ConfigMap with the new version 3. On failure, a `MigrationFailed` condition is set on the Deployment. The failed Job is kept for 60 seconds so its logs can be inspected, then replaced with a new one. +A running migration is never interrupted. If the image changes again while a Job's pod is running (a rollback, or two upgrades in a row), the operator waits for that Job to finish and then runs the migration for the new image. Aborting a non-transactional step such as Postgres's concurrent index build in migration 006 leaves an invalid index that the next run skips. To abort a migration that is stuck, delete the Job or set `migrationJob.activeDeadlineSeconds`. + The operator never changes the Deployment's replica count or pod template. On a new database, OpenFGA's readiness check (`MinimumSupportedDatastoreSchemaRevision`) keeps pods `NotReady` until the first migration has run. On an upgrade the existing schema already meets that minimum, so new pods serve on it while the Job applies the newer migrations, which is the same behaviour as the Helm hook flow. ## Prerequisites diff --git a/operator/internal/controller/migration_controller.go b/operator/internal/controller/migration_controller.go index 026ac66d..ba047cb9 100644 --- a/operator/internal/controller/migration_controller.go +++ b/operator/internal/controller/migration_controller.go @@ -88,17 +88,24 @@ func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) ( jobVersion := job.Annotations[AnnotationDesiredVersion] complete := isJobConditionTrue(job, batchv1.JobComplete) failedAt, failed := jobFailedAt(job) + running := !complete && !failed && ptr.Deref(job.Status.Ready, 0) > 0 outdated := jobVersion != desiredVersion // A Job whose pod cannot start (a bad secret reference, an image pull // error, an unschedulable pod) never fails on its own, so rebuild it once - // the Deployment's pod template has changed. A Job with a ready pod is left - // alone so a running migration is not cut off; one whose pod has just - // finished may still be rebuilt, which only re-runs a no-op migration. - if !outdated && !complete && !failed && job.Status.Active > 0 && ptr.Deref(job.Status.Ready, 0) == 0 { + // the Deployment's pod template has changed. A pod that has just finished + // may still be rebuilt, which only re-runs a no-op migration. + if !outdated && !complete && !failed && !running && job.Status.Active > 0 { want := r.buildMigrationJob(deployment, container, desiredVersion) outdated = job.Annotations[AnnotationPodTemplateHash] != want.Annotations[AnnotationPodTemplateHash] } if outdated { + // Never interrupt a running migration: a non-transactional step such as + // a concurrent index build that is aborted halfway leaves the schema in + // a state the next run does not repair. Replace the Job once it ends. + if running { + logger.V(1).Info("waiting for running migration job before replacing it", "job", job.Name, "jobVersion", jobVersion, "desiredVersion", desiredVersion) + return ctrl.Result{RequeueAfter: 10 * time.Second}, nil + } logger.Info("replacing migration job", "job", job.Name, "jobVersion", jobVersion, "desiredVersion", desiredVersion) if err := r.deleteJob(ctx, job); err != nil { return ctrl.Result{}, err diff --git a/operator/internal/controller/migration_controller_test.go b/operator/internal/controller/migration_controller_test.go index 6d19ecaa..34cc39d2 100644 --- a/operator/internal/controller/migration_controller_test.go +++ b/operator/internal/controller/migration_controller_test.go @@ -422,6 +422,50 @@ func TestReconcile_StartedJobWithOutdatedTemplate_Kept(t *testing.T) { } } +func TestReconcile_RunningJobForOtherVersion_KeptUntilItEnds(t *testing.T) { + // The image changes from v1.14.0 to v1.15.0 while the v1.14.0 migration is + // running. Interrupting it could leave the schema half-migrated, so the Job + // must be left alone and only replaced once it finishes. + old := newTestDeployment("openfga/openfga:v1.14.0") + job := newTestJob(old) + job.Status.Active = 1 + job.Status.Ready = ptr.To(int32(1)) + dep := newTestDeployment("openfga/openfga:v1.15.0") + r := newReconciler(t, nil, dep, job) + + if result := reconcileOnce(t, r); result.RequeueAfter != 10*time.Second { + t.Errorf("expected the running job to be polled, got %v", result.RequeueAfter) + } + kept, err := getJob(r) + if err != nil { + t.Fatalf("a running migration must not be deleted on a version change: %v", err) + } + if _, err := getStatus(r); !apierrors.IsNotFound(err) { + t.Errorf("no version may be recorded while the old job runs, got err=%v", err) + } + + kept.Status = batchv1.JobStatus{Succeeded: 1, Ready: ptr.To(int32(0)), Conditions: []batchv1.JobCondition{jobCondition(batchv1.JobComplete, time.Now())}} + if err := r.Status().Update(context.Background(), kept); err != nil { + t.Fatal(err) + } + reconcileOnce(t, r) + if _, err := getJob(r); !apierrors.IsNotFound(err) { + t.Fatalf("expected the finished v1.14.0 job to be replaced, got err=%v", err) + } + if _, err := getStatus(r); !apierrors.IsNotFound(err) { + t.Errorf("a v1.14.0 job must not be recorded as a v1.15.0 migration, got err=%v", err) + } + + reconcileOnce(t, r) + replacement, err := getJob(r) + if err != nil { + t.Fatalf("expected a migration job for the new version: %v", err) + } + if got := replacement.Annotations[AnnotationDesiredVersion]; got != "v1.15.0" { + t.Errorf("expected the new job to target v1.15.0, got %q", got) + } +} + // The legacy chart's Helm hook Job has the same name and carries the chart's // app.kubernetes.io/version label, which is the chart appVersion rather than // the image it ran. It must never be taken as proof of a migration. From ca25836f6255e56b03543103a6654e8e1d68ea7a Mon Sep 17 00:00:00 2001 From: Siddhant Khare <siddhant@usegitai.com> Date: Wed, 23 Sep 2026 16:24:56 +0530 Subject: [PATCH 62/70] fix(operator): preserve migration lifecycle inputs Track the complete migration Job identity, preserve migration-specific pod configuration, and prevent unsafe resource replacement. Harden retry cleanup and require parent chart release version propagation. --- .github/workflows/operator.yml | 47 ++- charts/openfga-operator/templates/role.yaml | 2 +- .../tests/watch_namespace_role_test.yaml | 14 + charts/openfga-operator/values.yaml | 3 +- charts/openfga/README.md | 4 +- charts/openfga/templates/NOTES.txt | 4 +- charts/openfga/templates/deployment.yaml | 33 ++ charts/openfga/tests/operator_mode_test.yaml | 56 ++++ charts/openfga/values.schema.json | 15 +- charts/openfga/values.yaml | 3 + docs/adr/002-operator-managed-migrations.md | 18 +- operator/README.md | 22 +- operator/internal/controller/helpers.go | 228 ++++++++++--- .../controller/migration_controller.go | 78 ++++- .../controller/migration_controller_test.go | 302 +++++++++++++++++- 15 files changed, 735 insertions(+), 94 deletions(-) diff --git a/.github/workflows/operator.yml b/.github/workflows/operator.yml index bc1f79d9..8bf36046 100644 --- a/.github/workflows/operator.yml +++ b/.github/workflows/operator.yml @@ -42,23 +42,52 @@ jobs: set -euo pipefail # CI publishes ghcr.io/openfga/openfga-operator:<appVersion> once and never # overwrites it, so a change to the image is only released when the - # operator chart's appVersion and version are bumped. + # operator chart's appVersion and version are bumped. The parent chart + # must also release the matching dependency version. image_inputs=(operator/cmd operator/internal operator/go.mod operator/go.sum operator/Dockerfile) if git diff --quiet "origin/$BASE...HEAD" -- "${image_inputs[@]}"; then echo "operator image inputs unchanged" exit 0 fi field() { grep "^$1:" "$2" | awk '{print $2}' | tr -d '"'; } - if ! git show "origin/$BASE:charts/openfga-operator/Chart.yaml" > /tmp/base-chart.yaml 2>/dev/null; then + dependency_version() { + awk ' + $1 == "-" && $2 == "name:" { + operator = ($3 == "openfga-operator") + next + } + operator && $1 == "version:" { + gsub(/"/, "", $2) + print $2 + exit + } + ' "$1" + } + + operator_version=$(field version charts/openfga-operator/Chart.yaml) + parent_dependency=$(dependency_version charts/openfga/Chart.yaml) + lock_dependency=$(dependency_version charts/openfga/Chart.lock) + if [[ "$parent_dependency" != "$operator_version" || "$lock_dependency" != "$operator_version" ]]; then + echo "::error file=charts/openfga/Chart.yaml::openfga-operator dependency and Chart.lock must match operator chart version ${operator_version}" + exit 1 + fi + + if git show "origin/$BASE:charts/openfga-operator/Chart.yaml" > /tmp/base-operator-chart.yaml 2>/dev/null; then + for f in appVersion version; do + if [[ "$(field $f /tmp/base-operator-chart.yaml)" == "$(field $f charts/openfga-operator/Chart.yaml)" ]]; then + echo "::error file=charts/openfga-operator/Chart.yaml::operator image inputs changed but $f is still $(field $f charts/openfga-operator/Chart.yaml); bump appVersion and version so the change is published" + exit 1 + fi + done + else echo "operator chart is new in this PR" - exit 0 fi - for f in appVersion version; do - if [[ "$(field $f /tmp/base-chart.yaml)" == "$(field $f charts/openfga-operator/Chart.yaml)" ]]; then - echo "::error file=charts/openfga-operator/Chart.yaml::operator image inputs changed but $f is still $(field $f charts/openfga-operator/Chart.yaml); bump appVersion and version so the change is published" - exit 1 - fi - done + + git show "origin/$BASE:charts/openfga/Chart.yaml" > /tmp/base-parent-chart.yaml + if [[ "$(field version /tmp/base-parent-chart.yaml)" == "$(field version charts/openfga/Chart.yaml)" ]]; then + echo "::error file=charts/openfga/Chart.yaml::operator image inputs changed but the parent chart version was not bumped" + exit 1 + fi - name: Set up Go uses: actions/setup-go@40f1582b2485089dde7abd97c1529aa768e1baff # v5.6.0 diff --git a/charts/openfga-operator/templates/role.yaml b/charts/openfga-operator/templates/role.yaml index e9e851f4..3ec0d99e 100644 --- a/charts/openfga-operator/templates/role.yaml +++ b/charts/openfga-operator/templates/role.yaml @@ -14,7 +14,7 @@ rules: verbs: ["patch"] - apiGroups: ["batch"] resources: ["jobs"] - verbs: ["get", "list", "watch", "create", "delete"] + verbs: ["get", "list", "watch", "create", "delete", "patch"] - apiGroups: [""] resources: ["configmaps"] verbs: ["get", "list", "watch", "create", "update"] diff --git a/charts/openfga-operator/tests/watch_namespace_role_test.yaml b/charts/openfga-operator/tests/watch_namespace_role_test.yaml index 687925d3..4cc98481 100644 --- a/charts/openfga-operator/tests/watch_namespace_role_test.yaml +++ b/charts/openfga-operator/tests/watch_namespace_role_test.yaml @@ -11,3 +11,17 @@ tests: - equal: path: metadata.namespace value: openfga-app + - contains: + path: rules + content: + apiGroups: + - batch + resources: + - jobs + verbs: + - get + - list + - watch + - create + - delete + - patch diff --git a/charts/openfga-operator/values.yaml b/charts/openfga-operator/values.yaml index 095b6efc..b6ba9f5c 100644 --- a/charts/openfga-operator/values.yaml +++ b/charts/openfga-operator/values.yaml @@ -66,7 +66,8 @@ migrationJob: # 0 disables the deadline. A migration that is cut off, such as an index build # on a large table, has to start over on the next attempt. activeDeadlineSeconds: 0 - # -- Seconds to keep completed/failed Job pods for log inspection before garbage collection. + # -- Seconds to keep completed Job pods for log inspection before garbage collection. + # Failed Jobs are retained for the operator's fixed 60-second retry delay. ttlSecondsAfterFinished: 300 resources: diff --git a/charts/openfga/README.md b/charts/openfga/README.md index f9ae8971..1cdef1e9 100644 --- a/charts/openfga/README.md +++ b/charts/openfga/README.md @@ -153,7 +153,7 @@ datastore: ### Running migrations with the operator -By default the chart runs database migrations from a Helm hook Job and gates the OpenFGA pods on it with an init container. Helm hooks are not run by Argo CD and conflict with `helm install --wait` and Flux, so the chart can instead install the [openfga-operator](../openfga-operator), which runs `openfga migrate` as a regular Job whenever the OpenFGA image changes: +By default the chart runs database migrations from a Helm hook Job and gates the OpenFGA pods on it with an init container. Helm hooks are not run by Argo CD and conflict with `helm install --wait` and Flux, so the chart can instead install the [openfga-operator](../openfga-operator), which runs `openfga migrate` as a regular Job whenever the OpenFGA image or migration inputs change: ```yaml openfga-operator: @@ -164,7 +164,7 @@ datastore: uriSecret: my-postgres-secret ``` -The operator only runs migrations; replicas, autoscaling and the pod template stay under the chart's control. It records the migrated version in the `<release>-migration-status` ConfigMap and sets a `MigrationFailed` condition on the Deployment if a migration fails. The migration Job runs as a dedicated `<release>-migration` service account (`migration.serviceAccount`), which can carry cloud IAM annotations for DDL permissions. Migrations only run when the image tag changes, so pin `image.tag` to a release rather than a floating tag. See the [operator README](../../operator/README.md) for how it works and its limitations. +The operator only runs migrations; replicas, autoscaling and the pod template stay under the chart's control. It records the migrated image and migration Job template identity in the `<release>-migration-status` ConfigMap and sets a `MigrationFailed` condition on the Deployment if a migration fails. The migration Job runs as a dedicated `<release>-migration` service account (`migration.serviceAccount`), which can carry cloud IAM annotations for DDL permissions. Migration-specific init containers, sidecars, volumes, mounts, resources, timeout, non-hook annotations, and labels are forwarded from `migrate.*` and `datastore.migrations.resources`. Change `migration.nonce` to rerun a migration after rotating referenced Secret data without changing the Secret name. See the [operator README](../../operator/README.md) for how it works and its limitations. ## Uninstalling the Chart diff --git a/charts/openfga/templates/NOTES.txt b/charts/openfga/templates/NOTES.txt index ed84cd71..2465fd7c 100644 --- a/charts/openfga/templates/NOTES.txt +++ b/charts/openfga/templates/NOTES.txt @@ -1,7 +1,7 @@ {{- if include "openfga.operatorMigrations" . }} NOTE: database migrations are run by the openfga-operator. Whenever the OpenFGA -image changes it runs the {{ include "openfga.fullname" . }}-migrate Job and records the migrated -version in the {{ include "openfga.fullname" . }}-migration-status ConfigMap. On a new database the +image or migration inputs change it runs the {{ include "openfga.fullname" . }}-migrate Job and +records the migrated identity in the {{ include "openfga.fullname" . }}-migration-status ConfigMap. On a new database the OpenFGA pods stay NotReady until the first migration completes. If the pods do not become ready, check the operator and the migration Job: diff --git a/charts/openfga/templates/deployment.yaml b/charts/openfga/templates/deployment.yaml index 6c81a287..30c3ec05 100644 --- a/charts/openfga/templates/deployment.yaml +++ b/charts/openfga/templates/deployment.yaml @@ -13,6 +13,39 @@ metadata: {{- if or .Values.migration.serviceAccount.create .Values.migration.serviceAccount.name }} openfga.dev/migration-service-account: '{{ include "openfga.migrationServiceAccountName" . }}' {{- end }} + {{- with .Values.migrate.extraInitContainers }} + openfga.dev/migration-init-containers: {{ . | toJson | quote }} + {{- end }} + {{- with .Values.migrate.sidecars }} + openfga.dev/migration-sidecars: {{ include "common.tplvalues.render" (dict "value" . "context" $) | fromYamlArray | toJson | quote }} + {{- end }} + {{- with .Values.migrate.extraVolumes }} + openfga.dev/migration-volumes: {{ . | toJson | quote }} + {{- end }} + {{- with .Values.migrate.extraVolumeMounts }} + openfga.dev/migration-volume-mounts: {{ . | toJson | quote }} + {{- end }} + {{- with .Values.datastore.migrations.resources }} + openfga.dev/migration-resources: {{ . | toJson | quote }} + {{- end }} + {{- with .Values.migrate.timeout }} + openfga.dev/migration-timeout: {{ . | quote }} + {{- end }} + {{- with .Values.migration.nonce }} + openfga.dev/migration-nonce: {{ . | quote }} + {{- end }} + {{- $migrationAnnotations := dict }} + {{- range $key, $value := .Values.migrate.annotations }} + {{- if not (hasPrefix "helm.sh/" $key) }} + {{- $_ := set $migrationAnnotations $key $value }} + {{- end }} + {{- end }} + {{- with $migrationAnnotations }} + openfga.dev/migration-annotations: {{ . | toJson | quote }} + {{- end }} + {{- with .Values.migrate.labels }} + openfga.dev/migration-labels: {{ . | toJson | quote }} + {{- end }} {{- end }} {{- with .Values.annotations }} {{- toYaml . | nindent 4 }} diff --git a/charts/openfga/tests/operator_mode_test.yaml b/charts/openfga/tests/operator_mode_test.yaml index 16dbc197..511a63db 100644 --- a/charts/openfga/tests/operator_mode_test.yaml +++ b/charts/openfga/tests/operator_mode_test.yaml @@ -18,6 +18,62 @@ tests: path: metadata.annotations["openfga.dev/migration-service-account"] value: RELEASE-NAME-openfga-migration + - it: should pass migration Job configuration to the operator + set: + openfga-operator.enabled: true + datastore.engine: postgres + datastore.migrations.resources.requests.cpu: 100m + migrate.timeout: 2m + migrate.extraInitContainers: + - name: prepare-proxy + image: busybox:1.36 + migrate.sidecars: + - name: database-proxy + image: "{{ .Release.Name }}-proxy:v2" + migrate.extraVolumes: + - name: proxy-config + secret: + secretName: database-proxy + migrate.extraVolumeMounts: + - name: proxy-config + mountPath: /credentials + readOnly: true + migrate.annotations: + helm.sh/hook: post-install + admission.example.com/inject: enabled + migrate.labels: + app.kubernetes.io/component: overridden + network-policy.example.com/database: allowed + migration.nonce: secret-rotation-2 + asserts: + - equal: + path: metadata.annotations["openfga.dev/migration-init-containers"] + value: '[{"image":"busybox:1.36","name":"prepare-proxy"}]' + - equal: + path: metadata.annotations["openfga.dev/migration-sidecars"] + value: '[{"image":"RELEASE-NAME-proxy:v2","name":"database-proxy"}]' + - equal: + path: metadata.annotations["openfga.dev/migration-volumes"] + value: '[{"name":"proxy-config","secret":{"secretName":"database-proxy"}}]' + - equal: + path: metadata.annotations["openfga.dev/migration-volume-mounts"] + value: '[{"mountPath":"/credentials","name":"proxy-config","readOnly":true}]' + - equal: + path: metadata.annotations["openfga.dev/migration-resources"] + value: '{"requests":{"cpu":"100m"}}' + - equal: + path: metadata.annotations["openfga.dev/migration-timeout"] + value: 2m + - equal: + path: metadata.annotations["openfga.dev/migration-nonce"] + value: secret-rotation-2 + - equal: + path: metadata.annotations["openfga.dev/migration-annotations"] + value: '{"admission.example.com/inject":"enabled"}' + - equal: + path: metadata.annotations["openfga.dev/migration-labels"] + value: '{"app.kubernetes.io/component":"overridden","network-policy.example.com/database":"allowed"}' + - it: should not set operator annotations when operator is disabled set: openfga-operator.enabled: false diff --git a/charts/openfga/values.schema.json b/charts/openfga/values.schema.json index bc56e79a..3b3b9fe0 100644 --- a/charts/openfga/values.schema.json +++ b/charts/openfga/values.schema.json @@ -1205,6 +1205,14 @@ }, "default": {} }, + "labels": { + "type": "object", + "description": "Map of labels to add to the migration job and pod", + "additionalProperties": { + "type": "string" + }, + "default": {} + }, "timeout": { "type": [ "string", @@ -1306,8 +1314,13 @@ }, "migration": { "type": "object", - "description": "Service account for the migration Jobs the operator creates. Only used when openfga-operator.enabled is true.", + "description": "Configuration for migration Jobs the operator creates. Only used when openfga-operator.enabled is true.", "properties": { + "nonce": { + "type": "string", + "description": "Arbitrary value included in the migration identity. Change it to rerun migrations when referenced Secret data changes without changing its name.", + "default": "" + }, "serviceAccount": { "type": "object", "properties": { diff --git a/charts/openfga/values.yaml b/charts/openfga/values.yaml index acef64f1..ee592aaa 100644 --- a/charts/openfga/values.yaml +++ b/charts/openfga/values.yaml @@ -408,6 +408,9 @@ openfga-operator: # -- Service account for the migration Jobs the operator creates. # Only used when openfga-operator.enabled is true. migration: + # -- Arbitrary value included in the migration identity. Change it to rerun + # migrations when referenced Secret data changes without changing its name. + nonce: "" serviceAccount: # -- Create a dedicated service account for migration Jobs. # The migration Job inherits env vars (including secretKeyRef) from the OpenFGA container. diff --git a/docs/adr/002-operator-managed-migrations.md b/docs/adr/002-operator-managed-migrations.md index e51923be..1072d6ee 100644 --- a/docs/adr/002-operator-managed-migrations.md +++ b/docs/adr/002-operator-managed-migrations.md @@ -77,10 +77,10 @@ The operator runs a **migration controller** that reconciles the OpenFGA Deploym ┌──────────────────────────────────────────────────────────┐ │ Operator Reconciliation │ │ │ -│ 1. Read Deployment → extract image tag (e.g. v1.14.0) │ +│ 1. Read Deployment and derive migration identity │ │ 2. Read ConfigMap/openfga-migration-status │ -│ └── "Last migrated version: v1.13.0" │ -│ 3. Versions differ → migration needed │ +│ └── "Last migrated image and pod template hash" │ +│ 3. Identities differ → migration needed │ │ 4. Create Job/openfga-migrate │ │ ├── ServiceAccount: openfga-migrator (DDL perms) │ │ ├── Image: openfga/openfga:v1.14.0 │ @@ -101,12 +101,12 @@ Readiness comes from OpenFGA itself: `IsReady()` reports `NOT_SERVING` while the **Rejected alternative — let the operator own the replica count:** the chart could omit `spec.replicas` (or render 0) and have the operator scale the Deployment up once the migration succeeds. Testing this showed three problems: switching an existing release to operator mode removes the field, so both Helm's three-way merge and server-side apply reset the Deployment to one replica until the migration finishes; `kubectl scale` and HPAs are overridden by the operator; and the scale-up buys nothing on upgrades, where the readiness check does not hold pods back. -#### Version tracking via ConfigMap +#### Migration identity tracking via ConfigMap -A ConfigMap (`openfga-migration-status`) records the last successfully migrated version. The operator compares this to the Deployment's image tag to determine if migration is needed. This is: +A ConfigMap (`openfga-migration-status`) records the last successfully migrated image version and migration Job pod template hash. The operator compares both values to the desired Job, so changes to datastore environment references, volumes, scheduling, init containers, sidecars, and other migration inputs trigger a new migration. Because Secret contents are not present in a Deployment, users can change `migration.nonce` to force a migration after rotating a referenced Secret in place. This is: - Simple to inspect (`kubectl get configmap openfga-migration-status -o yaml`) - Survives operator restarts -- Can be manually deleted to force re-migration (once the previous migration Job has been cleaned up) +- Can be manually deleted to force re-migration once the previous migration Job has been cleaned up #### Separate ServiceAccount for migrations @@ -206,6 +206,8 @@ Nothing is deleted outright — every change is gated on `openfga-operator.enabl | `values.yaml`: `openfga-operator.enabled` | Toggle the operator subchart | | `values.yaml`: `openfga-operator.migrationJob.*` | Migration Job backoff, deadline, and TTL configuration | | `values.yaml`: `migration.serviceAccount.*` | Separate ServiceAccount for migration Jobs | +| `values.yaml`: `migration.nonce` | Explicit rerun trigger for referenced Secret data changes | +| `values.yaml`: migration pod values | `migrate.extraInitContainers`, `migrate.sidecars`, volumes, mounts, resources, timeout, non-hook annotations, and labels are forwarded to operator Jobs | | `templates/serviceaccount.yaml`: second SA | Migration ServiceAccount | | `charts/openfga-operator/` | Operator subchart (conditional dependency) | @@ -230,5 +232,5 @@ Users on `openfga-operator.enabled: false` (the default) see identical rendered ### Risks - **Readiness relies on OpenFGA's schema check** — pods on a fresh database are held back only by `MinimumSupportedDatastoreSchemaRevision` in `pkg/storage/sqlcommon/sqlcommon.go`, and upgrades rely on each release working against the previous schema. Both are OpenFGA guarantees the Helm hook flow already depended on. -- **Migrations run as soon as the image changes** — as with the hook Job, nothing drains traffic first. Some migrations, such as MySQL's `008_collate_identifiers` in v1.18.0, block writes while tables are rebuilt; OpenFGA's runbook recommends draining traffic for those, which stays a manual step. -- **ConfigMap as state store** — if the ConfigMap is accidentally deleted, the operator records the version again from the completed Job while it exists, or re-runs the migration once it has been cleaned up (which is safe — `openfga migrate` is idempotent). +- **Migrations run as soon as their inputs change:** As with the hook Job, nothing drains traffic first. Some migrations, such as MySQL's `008_collate_identifiers` in v1.18.0, block writes while tables are rebuilt; OpenFGA's runbook recommends draining traffic for those, which stays a manual step. +- **ConfigMap as state store:** If the ConfigMap is accidentally deleted, the operator records the migration identity again from the completed Job while it exists, or reruns the migration once the Job has been cleaned up. `openfga migrate` is idempotent. diff --git a/operator/README.md b/operator/README.md index 9da24f71..8745eceb 100644 --- a/operator/README.md +++ b/operator/README.md @@ -1,13 +1,13 @@ # OpenFGA Operator -A Kubernetes operator that manages database migrations for OpenFGA deployments. Instead of relying on Helm hooks and init containers, the operator watches OpenFGA Deployments, detects version changes, and orchestrates migrations as regular Jobs. +A Kubernetes operator that manages database migrations for OpenFGA deployments. Instead of relying on Helm hooks and init containers, the operator watches OpenFGA Deployments, detects migration input changes, and orchestrates migrations as regular Jobs. This is **Stage 1** of the operator — focused solely on migration orchestration. See [ADR-001](../docs/adr/001-adopt-openfga-operator.md) for the full roadmap. ## How It Works 1. The operator watches Deployments in its configured namespace, which defaults to the operator pod's namespace, labeled `app.kubernetes.io/part-of: openfga` and `app.kubernetes.io/component: authorization-controller` -2. When a version change is detected (comparing the container image tag to the `{name}-migration-status` ConfigMap), the operator: +2. When the desired migration identity changes (comparing the rendered migration Job pod template to the `{name}-migration-status` ConfigMap), the operator: - Creates a migration Job running `openfga migrate` - Waits for the Job to complete - Updates the ConfigMap with the new version @@ -121,7 +121,7 @@ The operator accepts the following flags: | `--health-probe-bind-address` | `:8081` | Address the Kubernetes liveness and readiness probe endpoints bind to. Change only if the default port conflicts. | | `--backoff-limit` | `3` | Number of times a migration Job's pod can fail before the Job is considered failed. The operator then sets a `MigrationFailed` condition on the Deployment and replaces the Job 60 seconds after it failed. | | `--active-deadline-seconds` | `0` | Maximum wall-clock seconds a migration Job can run before Kubernetes terminates it. `0` means no deadline. A deadline cuts off long migrations, such as index builds or MySQL table rebuilds on large tables, which then start over on the next attempt. | -| `--ttl-seconds-after-finished` | `300` | Seconds Kubernetes keeps a completed or failed Job (and its pods) before garbage-collecting them, giving you time to inspect logs. | +| `--ttl-seconds-after-finished` | `300` | Seconds Kubernetes keeps a completed Job and its pod before garbage-collecting them. Failed Jobs do not receive a TTL and remain available for the operator's 60-second retry delay. | When deployed via the Helm subchart, these are configured through `values.yaml`. See `charts/openfga-operator/values.yaml` for all available options. @@ -134,11 +134,21 @@ The operator reads these annotations from the OpenFGA Deployment: | `openfga.dev/migration-enabled` | Must be `"true"` for the operator to manage migrations. Deployments without this annotation are ignored. Set by the Helm chart when `openfga-operator.enabled` and `datastore.applyMigrations` are true and the datastore is Postgres or MySQL. | | `openfga.dev/container-name` | The OpenFGA container in the pod spec. Defaults to `openfga`. | | `openfga.dev/migration-service-account` | The ServiceAccount to use for migration Jobs. Defaults to the Deployment's SA. | +| `openfga.dev/migration-init-containers` | JSON array of additional init containers for the migration Job. Generated from `migrate.extraInitContainers`. | +| `openfga.dev/migration-sidecars` | JSON array of additional containers for the migration Job. Generated from `migrate.sidecars`. | +| `openfga.dev/migration-volumes` | JSON array of additional volumes for the migration Job. Generated from `migrate.extraVolumes`. | +| `openfga.dev/migration-volume-mounts` | JSON array of additional mounts for the migration container. Generated from `migrate.extraVolumeMounts`. | +| `openfga.dev/migration-resources` | JSON resource requirements for the migration container. Generated from `datastore.migrations.resources`. | +| `openfga.dev/migration-timeout` | `OPENFGA_TIMEOUT` for the migration container. Generated from `migrate.timeout`. | +| `openfga.dev/migration-nonce` | Arbitrary value included in the migration identity. Generated from `migration.nonce`. | +| `openfga.dev/migration-annotations` | JSON map of non-Helm annotations for the migration Job and pod. Generated from `migrate.annotations`; `helm.sh/*` hook annotations are excluded. | +| `openfga.dev/migration-labels` | JSON map of additional labels for the migration Job and pod. Generated from `migrate.labels`; operator identity labels take precedence. | ## Limitations -- **Migrations key only on the image tag:** The operator compares the container image tag (or digest) to the `{name}-migration-status` ConfigMap. A mutable tag like `latest`, or a tag reused for a new build, is not seen as a change, so the migration is skipped — use immutable tags (e.g. `v1.14.0`) or pin by digest. A migration-needing change that keeps the same image — for example repointing `datastore.uri` at a different or restored database — also won't trigger a Job; delete the status ConfigMap (and the `{name}-migrate` Job, if it still exists) to run the migration again. -- **Legacy migration values:** `migrate.*` (extra volumes and mounts, init containers, sidecars, annotations, labels, timeout) and `datastore.migrations.resources` only apply to the legacy Helm hook Job. The operator's Job copies the OpenFGA container's image, env, volumes, resources, security context and scheduling instead, so put anything the migration needs (e.g. CA bundles) in the top-level `extraVolumes`, `extraVolumeMounts` and `extraEnvVars`. -- **Single-container migration Job:** The Job runs one container (`openfga migrate`) and injects no sidecars or extra init containers, so databases reached through a sidecar proxy (Cloud SQL Auth Proxy, AlloyDB) aren't supported for operator-managed migrations. A sidecar injected into every pod in the namespace that doesn't exit on its own (e.g. an Istio sidecar) keeps the Job pod running and stops the Job from completing. +- **Secret contents are not observable:** The migration identity covers the image, environment references, pod configuration, migration-specific containers, and `migration.nonce`. Kubernetes does not expose referenced Secret contents through the Deployment, so change `migration.nonce` when rotating a Secret in place and a migration must rerun. +- **Mutable image contents are not observable:** Reusing a tag such as `latest` does not change the Deployment's image reference. Use immutable tags or digests, or change `migration.nonce` when deliberately replacing the contents of a mutable tag. +- **Helm hook metadata:** Operator-managed Jobs ignore `helm.sh/*` entries in `migrate.annotations`. Other migration annotations and labels are forwarded, but cannot override the operator's identity labels. +- **Sidecar completion:** Containers configured through `migrate.sidecars` must exit after the migration completes. A sidecar that runs indefinitely keeps the Job pod running and prevents the Job from completing. - **Job pod labels:** The migration pod is labelled `app.kubernetes.io/part-of: openfga` and `app.kubernetes.io/component: migration`, not with the OpenFGA Deployment's `app.kubernetes.io/name`/`instance` labels (which would make it a Service endpoint). A NetworkPolicy that allows database egress only for the OpenFGA pods' labels needs a rule for the migration pod too. - **One namespace per operator:** The operator reconciles every opted-in OpenFGA Deployment in its watch namespace. Operators installed by several releases in one namespace share a leader election lease, so only one of them is active at a time. diff --git a/operator/internal/controller/helpers.go b/operator/internal/controller/helpers.go index a82f1799..3b0baf9d 100644 --- a/operator/internal/controller/helpers.go +++ b/operator/internal/controller/helpers.go @@ -5,6 +5,7 @@ import ( "crypto/sha256" "encoding/json" "fmt" + "io" "strings" "time" @@ -33,6 +34,15 @@ const ( AnnotationMigrationEnabled = "openfga.dev/migration-enabled" AnnotationContainerName = "openfga.dev/container-name" AnnotationMigrationServiceAccount = "openfga.dev/migration-service-account" + AnnotationMigrationInitContainers = "openfga.dev/migration-init-containers" + AnnotationMigrationSidecars = "openfga.dev/migration-sidecars" + AnnotationMigrationVolumes = "openfga.dev/migration-volumes" + AnnotationMigrationVolumeMounts = "openfga.dev/migration-volume-mounts" + AnnotationMigrationResources = "openfga.dev/migration-resources" + AnnotationMigrationTimeout = "openfga.dev/migration-timeout" + AnnotationMigrationNonce = "openfga.dev/migration-nonce" + AnnotationMigrationAnnotations = "openfga.dev/migration-annotations" + AnnotationMigrationLabels = "openfga.dev/migration-labels" // Annotations set on migration Jobs: the version the Job migrates to, and a // hash of the pod template it was built from. @@ -103,57 +113,126 @@ func ownerReference(deployment *appsv1.Deployment) metav1.OwnerReference { } } +func isOperatorManagedResourceForDeployment(obj metav1.Object, deployment *appsv1.Deployment) bool { + if obj.GetLabels()[LabelManagedBy] != LabelManagedByValue { + return false + } + for _, owner := range obj.GetOwnerReferences() { + if ptr.Deref(owner.Controller, false) && + owner.APIVersion == "apps/v1" && + owner.Kind == "Deployment" && + owner.Name == deployment.Name { + return true + } + } + return false +} + +func isLegacyMigrationJob(job *batchv1.Job) bool { + return job.Labels[LabelManagedBy] == "Helm" && job.Annotations["helm.sh/hook"] != "" +} + // buildMigrationJob constructs a Job that runs "openfga migrate" with the // OpenFGA container's image, environment, volumes and scheduling. -func (r *MigrationReconciler) buildMigrationJob(deployment *appsv1.Deployment, container *corev1.Container, version string) *batchv1.Job { +func (r *MigrationReconciler) buildMigrationJob(deployment *appsv1.Deployment, container *corev1.Container, version string) (*batchv1.Job, error) { podSpec := deployment.Spec.Template.Spec serviceAccount := deployment.Annotations[AnnotationMigrationServiceAccount] if serviceAccount == "" { serviceAccount = podSpec.ServiceAccountName } + initContainers, err := annotationJSON[[]corev1.Container](deployment, AnnotationMigrationInitContainers) + if err != nil { + return nil, err + } + sidecars, err := annotationJSON[[]corev1.Container](deployment, AnnotationMigrationSidecars) + if err != nil { + return nil, err + } + extraVolumes, err := annotationJSON[[]corev1.Volume](deployment, AnnotationMigrationVolumes) + if err != nil { + return nil, err + } + extraVolumeMounts, err := annotationJSON[[]corev1.VolumeMount](deployment, AnnotationMigrationVolumeMounts) + if err != nil { + return nil, err + } + migrationAnnotations, err := annotationJSON[map[string]string](deployment, AnnotationMigrationAnnotations) + if err != nil { + return nil, err + } + migrationLabels, err := annotationJSON[map[string]string](deployment, AnnotationMigrationLabels) + if err != nil { + return nil, err + } + + resources := container.Resources + if deployment.Annotations[AnnotationMigrationResources] != "" { + resources, err = annotationJSON[corev1.ResourceRequirements](deployment, AnnotationMigrationResources) + if err != nil { + return nil, err + } + } + env := append([]corev1.EnvVar(nil), container.Env...) + if timeout := deployment.Annotations[AnnotationMigrationTimeout]; timeout != "" && !hasEnvVar(env, "OPENFGA_TIMEOUT") { + env = append(env, corev1.EnvVar{Name: "OPENFGA_TIMEOUT", Value: timeout}) + } + + podAnnotations := mergeStringMaps(migrationAnnotations, nil) + delete(podAnnotations, AnnotationMigrationNonce) + if nonce := deployment.Annotations[AnnotationMigrationNonce]; nonce != "" { + podAnnotations[AnnotationMigrationNonce] = nonce + } + jobLabels := mergeStringMaps(migrationLabels, map[string]string{ + LabelPartOf: LabelPartOfValue, + LabelComponent: "migration", + LabelManagedBy: LabelManagedByValue, + }) + podLabels := mergeStringMaps(migrationLabels, map[string]string{ + LabelPartOf: LabelPartOfValue, + LabelComponent: "migration", + }) + jobAnnotations := mergeStringMaps(migrationAnnotations, map[string]string{ + AnnotationDesiredVersion: version, + }) + containers := append([]corev1.Container{{ + Name: "migrate-database", + Image: container.Image, + ImagePullPolicy: container.ImagePullPolicy, + Args: []string{"migrate"}, + Env: env, + EnvFrom: container.EnvFrom, + Resources: resources, + VolumeMounts: mergeVolumeMounts(container.VolumeMounts, extraVolumeMounts), + SecurityContext: container.SecurityContext, + }}, sidecars...) + job := &batchv1.Job{ ObjectMeta: metav1.ObjectMeta{ - Name: migrationJobName(deployment.Name), - Namespace: deployment.Namespace, - Labels: map[string]string{ - LabelPartOf: LabelPartOfValue, - LabelComponent: "migration", - LabelManagedBy: LabelManagedByValue, - }, - Annotations: map[string]string{AnnotationDesiredVersion: version}, + Name: migrationJobName(deployment.Name), + Namespace: deployment.Namespace, + Labels: jobLabels, + Annotations: jobAnnotations, OwnerReferences: []metav1.OwnerReference{ownerReference(deployment)}, }, Spec: batchv1.JobSpec{ - BackoffLimit: ptr.To(r.BackoffLimit), - TTLSecondsAfterFinished: ptr.To(r.TTLSecondsAfterFinished), + BackoffLimit: ptr.To(r.BackoffLimit), Template: corev1.PodTemplateSpec{ ObjectMeta: metav1.ObjectMeta{ - Labels: map[string]string{ - LabelPartOf: LabelPartOfValue, - LabelComponent: "migration", - }, + Labels: podLabels, + Annotations: podAnnotations, }, Spec: corev1.PodSpec{ ServiceAccountName: serviceAccount, RestartPolicy: corev1.RestartPolicyNever, ImagePullSecrets: podSpec.ImagePullSecrets, SecurityContext: podSpec.SecurityContext, - Containers: []corev1.Container{{ - Name: "migrate-database", - Image: container.Image, - ImagePullPolicy: container.ImagePullPolicy, - Args: []string{"migrate"}, - Env: container.Env, - EnvFrom: container.EnvFrom, - Resources: container.Resources, - VolumeMounts: container.VolumeMounts, - SecurityContext: container.SecurityContext, - }}, - Volumes: podSpec.Volumes, - NodeSelector: podSpec.NodeSelector, - Tolerations: podSpec.Tolerations, - Affinity: podSpec.Affinity, + InitContainers: initContainers, + Containers: containers, + Volumes: mergeVolumes(podSpec.Volumes, extraVolumes), + NodeSelector: podSpec.NodeSelector, + Tolerations: podSpec.Tolerations, + Affinity: podSpec.Affinity, }, }, }, @@ -162,7 +241,78 @@ func (r *MigrationReconciler) buildMigrationJob(deployment *appsv1.Deployment, c job.Spec.ActiveDeadlineSeconds = ptr.To(r.ActiveDeadlineSeconds) } job.Annotations[AnnotationPodTemplateHash] = podTemplateHash(&job.Spec.Template) - return job + return job, nil +} + +func annotationJSON[T any](deployment *appsv1.Deployment, annotation string) (T, error) { + var value T + raw := deployment.Annotations[annotation] + if raw == "" { + return value, nil + } + decoder := json.NewDecoder(strings.NewReader(raw)) + decoder.DisallowUnknownFields() + if err := decoder.Decode(&value); err != nil { + return value, fmt.Errorf("decoding %s annotation on deployment %s/%s: %w", annotation, deployment.Namespace, deployment.Name, err) + } + if err := decoder.Decode(&struct{}{}); err != io.EOF { + return value, fmt.Errorf("decoding %s annotation on deployment %s/%s: trailing JSON data", annotation, deployment.Namespace, deployment.Name) + } + return value, nil +} + +func hasEnvVar(env []corev1.EnvVar, name string) bool { + for i := range env { + if env[i].Name == name { + return true + } + } + return false +} + +func mergeVolumes(base, extra []corev1.Volume) []corev1.Volume { + merged := append([]corev1.Volume(nil), base...) + index := make(map[string]int, len(merged)) + for i := range merged { + index[merged[i].Name] = i + } + for _, volume := range extra { + if i, ok := index[volume.Name]; ok { + merged[i] = volume + continue + } + index[volume.Name] = len(merged) + merged = append(merged, volume) + } + return merged +} + +func mergeVolumeMounts(base, extra []corev1.VolumeMount) []corev1.VolumeMount { + merged := append([]corev1.VolumeMount(nil), base...) + index := make(map[string]int, len(merged)) + for i := range merged { + index[merged[i].MountPath] = i + } + for _, mount := range extra { + if i, ok := index[mount.MountPath]; ok { + merged[i] = mount + continue + } + index[mount.MountPath] = len(merged) + merged = append(merged, mount) + } + return merged +} + +func mergeStringMaps(base, overrides map[string]string) map[string]string { + merged := make(map[string]string, len(base)+len(overrides)) + for key, value := range base { + merged[key] = value + } + for key, value := range overrides { + merged[key] = value + } + return merged } func podTemplateHash(template *corev1.PodTemplateSpec) string { @@ -173,13 +323,18 @@ func podTemplateHash(template *corev1.PodTemplateSpec) string { return fmt.Sprintf("%x", sha256.Sum256(b))[:16] } -// updateMigrationStatus records the migrated version in the status ConfigMap. -func updateMigrationStatus(ctx context.Context, c client.Client, deployment *appsv1.Deployment, version, jobName string) error { +// updateMigrationStatus records the migrated version and Job template identity. +func updateMigrationStatus(ctx context.Context, c client.Client, deployment *appsv1.Deployment, version, podTemplateHash, jobName string) error { cm := &corev1.ConfigMap{ObjectMeta: metav1.ObjectMeta{ Name: migrationConfigMapName(deployment.Name), Namespace: deployment.Namespace, }} _, err := controllerutil.CreateOrUpdate(ctx, c, cm, func() error { + if cm.ResourceVersion != "" && + !metav1.IsControlledBy(cm, deployment) && + !isOperatorManagedResourceForDeployment(cm, deployment) { + return fmt.Errorf("ConfigMap %s/%s already exists and is not managed by the OpenFGA operator", cm.Namespace, cm.Name) + } cm.Labels = map[string]string{ LabelPartOf: LabelPartOfValue, LabelComponent: "migration", @@ -188,9 +343,10 @@ func updateMigrationStatus(ctx context.Context, c client.Client, deployment *app // Reset on every write in case the Deployment was recreated with a new UID. cm.OwnerReferences = []metav1.OwnerReference{ownerReference(deployment)} cm.Data = map[string]string{ - "version": version, - "migratedAt": time.Now().UTC().Format(time.RFC3339), - "jobName": jobName, + "version": version, + "podTemplateHash": podTemplateHash, + "migratedAt": time.Now().UTC().Format(time.RFC3339), + "jobName": jobName, } return nil }) diff --git a/operator/internal/controller/migration_controller.go b/operator/internal/controller/migration_controller.go index 026ac66d..efae54ea 100644 --- a/operator/internal/controller/migration_controller.go +++ b/operator/internal/controller/migration_controller.go @@ -23,7 +23,7 @@ import ( const retryDelay = 60 * time.Second // MigrationReconciler watches OpenFGA Deployments and runs a database -// migration Job whenever the OpenFGA image version changes. +// migration Job whenever its image or migration inputs change. type MigrationReconciler struct { client.Client @@ -52,14 +52,24 @@ func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) ( return ctrl.Result{}, err } desiredVersion := extractImageTag(container.Image) + desiredJob, err := r.buildMigrationJob(deployment, container, desiredVersion) + if err != nil { + return ctrl.Result{}, err + } + desiredPodTemplateHash := desiredJob.Annotations[AnnotationPodTemplateHash] status := &corev1.ConfigMap{} err = r.Get(ctx, types.NamespacedName{Name: migrationConfigMapName(req.Name), Namespace: req.Namespace}, status) if err != nil && !apierrors.IsNotFound(err) { return ctrl.Result{}, fmt.Errorf("getting migration status: %w", err) } + statusOwnedByDeployment := err == nil && metav1.IsControlledBy(status, deployment) + if err == nil && !statusOwnedByDeployment && !isOperatorManagedResourceForDeployment(status, deployment) { + return ctrl.Result{}, fmt.Errorf("migration status ConfigMap %s/%s already exists and is not managed by this Deployment", status.Namespace, status.Name) + } currentVersion := status.Data["version"] - if currentVersion == desiredVersion { + currentPodTemplateHash := status.Data["podTemplateHash"] + if statusOwnedByDeployment && currentVersion == desiredVersion && currentPodTemplateHash == desiredPodTemplateHash { _, err := r.patchCondition(ctx, deployment, clearMigrationFailedCondition) return ctrl.Result{}, err } @@ -67,7 +77,7 @@ func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) ( job := &batchv1.Job{} err = r.Get(ctx, types.NamespacedName{Name: migrationJobName(req.Name), Namespace: req.Namespace}, job) if apierrors.IsNotFound(err) { - job = r.buildMigrationJob(deployment, container, desiredVersion) + job = desiredJob if err := r.Create(ctx, job); err != nil { if apierrors.IsAlreadyExists(err) { // The cache has not caught up with a Job created by an earlier reconcile. @@ -81,24 +91,27 @@ func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) ( if err != nil { return ctrl.Result{}, fmt.Errorf("getting migration job: %w", err) } + jobOwnedByDeployment := metav1.IsControlledBy(job, deployment) + replaceableJob := jobOwnedByDeployment || isOperatorManagedResourceForDeployment(job, deployment) || isLegacyMigrationJob(job) + if !replaceableJob { + return ctrl.Result{}, fmt.Errorf("migration Job %s/%s already exists and is not managed by this Deployment", job.Namespace, job.Name) + } - // Only a Job this operator created for the desired version is trusted. + // Only a Job this operator created for the desired migration inputs is trusted. // Anything else under the same name, such as a Job for a previous image or // the chart's legacy Helm hook Job, is replaced. jobVersion := job.Annotations[AnnotationDesiredVersion] + jobPodTemplateHash := job.Annotations[AnnotationPodTemplateHash] complete := isJobConditionTrue(job, batchv1.JobComplete) failedAt, failed := jobFailedAt(job) - outdated := jobVersion != desiredVersion - // A Job whose pod cannot start (a bad secret reference, an image pull - // error, an unschedulable pod) never fails on its own, so rebuild it once - // the Deployment's pod template has changed. A Job with a ready pod is left - // alone so a running migration is not cut off; one whose pod has just - // finished may still be rebuilt, which only re-runs a no-op migration. - if !outdated && !complete && !failed && job.Status.Active > 0 && ptr.Deref(job.Status.Ready, 0) == 0 { - want := r.buildMigrationJob(deployment, container, desiredVersion) - outdated = job.Annotations[AnnotationPodTemplateHash] != want.Annotations[AnnotationPodTemplateHash] - } + outdated := !jobOwnedByDeployment || jobVersion != desiredVersion || jobPodTemplateHash != desiredPodTemplateHash if outdated { + if !complete && !failed && (job.Status.Ready != nil && *job.Status.Ready > 0 || job.Status.Succeeded > 0) { + // Never interrupt a migration that has started. Once it reaches a + // terminal state, the next reconciliation replaces it with a Job + // for the latest desired inputs. + return ctrl.Result{RequeueAfter: 10 * time.Second}, nil + } logger.Info("replacing migration job", "job", job.Name, "jobVersion", jobVersion, "desiredVersion", desiredVersion) if err := r.deleteJob(ctx, job); err != nil { return ctrl.Result{}, err @@ -107,7 +120,10 @@ func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) ( } if complete { - if err := updateMigrationStatus(ctx, r.Client, deployment, desiredVersion, job.Name); err != nil { + if err := r.setCompletedJobTTL(ctx, job); err != nil { + return ctrl.Result{}, err + } + if err := updateMigrationStatus(ctx, r.Client, deployment, desiredVersion, desiredPodTemplateHash, job.Name); err != nil { return ctrl.Result{}, err } logger.Info("migration succeeded", "version", desiredVersion) @@ -140,8 +156,36 @@ func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) ( return ctrl.Result{RequeueAfter: 10 * time.Second}, nil } +func (r *MigrationReconciler) setCompletedJobTTL(ctx context.Context, job *batchv1.Job) error { + if job.Spec.TTLSecondsAfterFinished != nil && *job.Spec.TTLSecondsAfterFinished == r.TTLSecondsAfterFinished { + return nil + } + patch := client.MergeFromWithOptions(job.DeepCopy(), client.MergeFromWithOptimisticLock{}) + job.Spec.TTLSecondsAfterFinished = ptr.To(r.TTLSecondsAfterFinished) + if err := r.Patch(ctx, job, patch); err != nil { + return fmt.Errorf("setting completed migration Job TTL: %w", err) + } + return nil +} + func (r *MigrationReconciler) deleteJob(ctx context.Context, job *batchv1.Job) error { - err := r.Delete(ctx, job, client.PropagationPolicy(metav1.DeletePropagationBackground)) + options := []client.DeleteOption{client.PropagationPolicy(metav1.DeletePropagationForeground)} + preconditions := client.Preconditions{} + hasPreconditions := false + if job.UID != "" { + uid := job.UID + preconditions.UID = &uid + hasPreconditions = true + } + if job.ResourceVersion != "" { + resourceVersion := job.ResourceVersion + preconditions.ResourceVersion = &resourceVersion + hasPreconditions = true + } + if hasPreconditions { + options = append(options, preconditions) + } + err := r.Delete(ctx, job, options...) if client.IgnoreNotFound(err) != nil { return fmt.Errorf("deleting migration job %s: %w", job.Name, err) } @@ -153,7 +197,7 @@ func (r *MigrationReconciler) deleteJob(ctx context.Context, job *batchv1.Job) e // merges conditions by type, so the Deployment controller's own conditions are // left alone. func (r *MigrationReconciler) patchCondition(ctx context.Context, deployment *appsv1.Deployment, update func(*appsv1.Deployment) bool) (bool, error) { - patch := client.StrategicMergeFrom(deployment.DeepCopy()) + patch := client.StrategicMergeFrom(deployment.DeepCopy(), client.MergeFromWithOptimisticLock{}) if !update(deployment) { return false, nil } diff --git a/operator/internal/controller/migration_controller_test.go b/operator/internal/controller/migration_controller_test.go index 6d19ecaa..0dc675fa 100644 --- a/operator/internal/controller/migration_controller_test.go +++ b/operator/internal/controller/migration_controller_test.go @@ -10,6 +10,7 @@ import ( batchv1 "k8s.io/api/batch/v1" corev1 "k8s.io/api/core/v1" apierrors "k8s.io/apimachinery/pkg/api/errors" + "k8s.io/apimachinery/pkg/api/resource" metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" "k8s.io/apimachinery/pkg/runtime" "k8s.io/apimachinery/pkg/types" @@ -66,7 +67,10 @@ func newTestDeployment(image string) *appsv1.Deployment { // newTestJob returns the migration Job the operator builds for dep. func newTestJob(dep *appsv1.Deployment, conditions ...batchv1.JobCondition) *batchv1.Job { container := &dep.Spec.Template.Spec.Containers[0] - job := (&MigrationReconciler{}).buildMigrationJob(dep, container, extractImageTag(container.Image)) + job, err := (&MigrationReconciler{}).buildMigrationJob(dep, container, extractImageTag(container.Image)) + if err != nil { + panic(err) + } job.Status.Conditions = conditions return job } @@ -75,10 +79,21 @@ func jobCondition(t batchv1.JobConditionType, at time.Time) batchv1.JobCondition return batchv1.JobCondition{Type: t, Status: corev1.ConditionTrue, LastTransitionTime: metav1.NewTime(at)} } -func newStatus(version string) *corev1.ConfigMap { +func newStatus(dep *appsv1.Deployment) *corev1.ConfigMap { + job := newTestJob(dep) return &corev1.ConfigMap{ - ObjectMeta: metav1.ObjectMeta{Name: statusKey.Name, Namespace: statusKey.Namespace}, - Data: map[string]string{"version": version}, + ObjectMeta: metav1.ObjectMeta{ + Name: statusKey.Name, + Namespace: statusKey.Namespace, + Labels: map[string]string{LabelManagedBy: LabelManagedByValue}, + OwnerReferences: []metav1.OwnerReference{ + ownerReference(dep), + }, + }, + Data: map[string]string{ + "version": job.Annotations[AnnotationDesiredVersion], + "podTemplateHash": job.Annotations[AnnotationPodTemplateHash], + }, } } @@ -181,6 +196,9 @@ func TestReconcile_FirstInstall_CreatesJob(t *testing.T) { if job.Spec.ActiveDeadlineSeconds != nil { t.Errorf("expected no deadline by default, got %d", *job.Spec.ActiveDeadlineSeconds) } + if job.Spec.TTLSecondsAfterFinished != nil { + t.Errorf("expected TTL to remain unset until the Job succeeds, got %d", *job.Spec.TTLSecondsAfterFinished) + } if len(job.OwnerReferences) != 1 || !ptr.Deref(job.OwnerReferences[0].Controller, false) || job.OwnerReferences[0].BlockOwnerDeletion != nil { t.Errorf("expected a single controller owner reference without blockOwnerDeletion, got %+v", job.OwnerReferences) } @@ -221,7 +239,7 @@ func TestReconcile_JobAlreadyExistsOnCreate_Requeues(t *testing.T) { func TestReconcile_VersionMatch_ClearsFailedCondition(t *testing.T) { dep := newTestDeployment("openfga/openfga:v1.14.0") dep.Status.Conditions = []appsv1.DeploymentCondition{{Type: "MigrationFailed", Status: corev1.ConditionTrue}} - r := newReconciler(t, nil, dep, newStatus("v1.14.0")) + r := newReconciler(t, nil, dep, newStatus(dep)) if result := reconcileOnce(t, r); result.RequeueAfter != 0 { t.Errorf("expected no requeue when versions match, got %v", result.RequeueAfter) @@ -241,7 +259,7 @@ func TestReconcile_VersionMatch_ClearsFailedCondition(t *testing.T) { func TestReconcile_VersionMatch_StatusPatchError(t *testing.T) { dep := newTestDeployment("openfga/openfga:v1.14.0") dep.Status.Conditions = []appsv1.DeploymentCondition{{Type: "MigrationFailed", Status: corev1.ConditionTrue}} - r := newReconciler(t, failStatusPatch, dep, newStatus("v1.14.0")) + r := newReconciler(t, failStatusPatch, dep, newStatus(dep)) if _, err := r.Reconcile(context.Background(), ctrl.Request{NamespacedName: deploymentKey}); err == nil { t.Fatal("expected the status patch error to be returned") @@ -251,7 +269,8 @@ func TestReconcile_VersionMatch_StatusPatchError(t *testing.T) { func TestReconcile_JobSucceeded_CreatesStatus(t *testing.T) { dep := newTestDeployment("openfga/openfga:v1.14.0") dep.Status.Conditions = []appsv1.DeploymentCondition{{Type: "MigrationFailed", Status: corev1.ConditionTrue, Reason: "MigrationJobFailed"}} - r := newReconciler(t, nil, dep, newTestJob(dep, jobCondition(batchv1.JobComplete, time.Now()))) + job := newTestJob(dep, jobCondition(batchv1.JobComplete, time.Now())) + r := newReconciler(t, nil, dep, job) if result := reconcileOnce(t, r); result.RequeueAfter != 0 { t.Errorf("expected no requeue after success, got %v", result.RequeueAfter) @@ -260,23 +279,48 @@ func TestReconcile_JobSucceeded_CreatesStatus(t *testing.T) { if err != nil { t.Fatalf("expected migration status ConfigMap: %v", err) } - if cm.Data["version"] != "v1.14.0" || cm.Data["jobName"] != jobKey.Name { + if cm.Data["version"] != "v1.14.0" || cm.Data["podTemplateHash"] != job.Annotations[AnnotationPodTemplateHash] || cm.Data["jobName"] != jobKey.Name { t.Errorf("unexpected status data: %v", cm.Data) } if len(cm.OwnerReferences) != 1 || cm.OwnerReferences[0].UID != "test-uid-123" { t.Errorf("expected the Deployment to own the status ConfigMap, got %+v", cm.OwnerReferences) } + completedJob, err := getJob(r) + if err != nil { + t.Fatalf("expected the completed Job: %v", err) + } + if got := ptr.Deref(completedJob.Spec.TTLSecondsAfterFinished, -1); got != DefaultTTLSecondsAfterFinished { + t.Errorf("expected completed Job TTL %d, got %d", DefaultTTLSecondsAfterFinished, got) + } cond := findCondition(getDeployment(t, r).Status.Conditions, "MigrationFailed") if cond == nil || cond.Status != corev1.ConditionFalse || cond.Reason != "MigrationSucceeded" { t.Errorf("expected MigrationFailed=False/MigrationSucceeded, got %+v", cond) } } +func TestReconcile_JobSucceeded_AppliesZeroTTL(t *testing.T) { + dep := newTestDeployment("openfga/openfga:v1.14.0") + r := newReconciler(t, nil, dep, newTestJob(dep, jobCondition(batchv1.JobComplete, time.Now()))) + r.TTLSecondsAfterFinished = 0 + + reconcileOnce(t, r) + + job, err := getJob(r) + if err != nil { + t.Fatalf("expected the fake client to retain the completed Job: %v", err) + } + if job.Spec.TTLSecondsAfterFinished == nil || *job.Spec.TTLSecondsAfterFinished != 0 { + t.Errorf("expected completed Job TTL 0, got %v", job.Spec.TTLSecondsAfterFinished) + } +} + func TestReconcile_JobSucceeded_UpdatesStatus(t *testing.T) { - status := newStatus("v1.13.0") - status.OwnerReferences = []metav1.OwnerReference{{APIVersion: "apps/v1", Kind: "Deployment", Name: "openfga", UID: "old-uid"}} + oldDep := newTestDeployment("openfga/openfga:v1.13.0") + oldDep.UID = "old-uid" + status := newStatus(oldDep) dep := newTestDeployment("openfga/openfga:v1.14.0") - r := newReconciler(t, nil, dep, status, newTestJob(dep, jobCondition(batchv1.JobComplete, time.Now()))) + job := newTestJob(dep, jobCondition(batchv1.JobComplete, time.Now())) + r := newReconciler(t, nil, dep, status, job) reconcileOnce(t, r) @@ -287,6 +331,9 @@ func TestReconcile_JobSucceeded_UpdatesStatus(t *testing.T) { if cm.Data["version"] != "v1.14.0" { t.Errorf("expected version v1.14.0, got %q", cm.Data["version"]) } + if cm.Data["podTemplateHash"] != job.Annotations[AnnotationPodTemplateHash] { + t.Errorf("expected the completed Job's pod template hash, got %q", cm.Data["podTemplateHash"]) + } if cm.OwnerReferences[0].UID != "test-uid-123" { t.Errorf("expected owner reference to be reset to the current Deployment, got %+v", cm.OwnerReferences) } @@ -315,6 +362,13 @@ func TestReconcile_JobFailed_KeepsJobUntilRetryDelay(t *testing.T) { if _, err := getJob(r); err != nil { t.Errorf("expected the failed job to be kept during the retry delay: %v", err) } + job, err := getJob(r) + if err != nil { + t.Fatalf("expected the failed Job: %v", err) + } + if job.Spec.TTLSecondsAfterFinished != nil { + t.Errorf("failed Jobs must not have a completion TTL, got %d", *job.Spec.TTLSecondsAfterFinished) + } cond := findCondition(getDeployment(t, r).Status.Conditions, "MigrationFailed") if cond == nil || cond.Status != corev1.ConditionTrue || cond.Reason != "MigrationJobFailed" { t.Errorf("expected MigrationFailed=True, got %+v", cond) @@ -397,6 +451,91 @@ func TestReconcile_PendingJobWithOutdatedTemplate_Replaced(t *testing.T) { } } +func TestReconcile_StatusWithOutdatedMigrationInputs_Reruns(t *testing.T) { + tests := []struct { + name string + change func(*appsv1.Deployment) + }{ + { + name: "datastore URI", + change: func(dep *appsv1.Deployment) { + dep.Spec.Template.Spec.Containers[0].Env[1].Value = "postgres://other.example.com/openfga" + }, + }, + { + name: "migration nonce", + change: func(dep *appsv1.Deployment) { + dep.Annotations[AnnotationMigrationNonce] = "secret-rotation-2" + }, + }, + } + for _, tt := range tests { + t.Run(tt.name, func(t *testing.T) { + dep := newTestDeployment("openfga/openfga:v1.14.0") + status := newStatus(dep) + tt.change(dep) + r := newReconciler(t, nil, dep, status) + + reconcileOnce(t, r) + job, err := getJob(r) + if err != nil { + t.Fatalf("expected changed migration inputs to create a Job: %v", err) + } + if job.Annotations[AnnotationPodTemplateHash] == status.Data["podTemplateHash"] { + t.Error("expected changed migration inputs to produce a new identity") + } + }) + } +} + +func TestReconcile_VersionOnlyStatus_Reruns(t *testing.T) { + dep := newTestDeployment("openfga/openfga:v1.14.0") + status := &corev1.ConfigMap{ + ObjectMeta: metav1.ObjectMeta{ + Name: statusKey.Name, + Namespace: statusKey.Namespace, + Labels: map[string]string{LabelManagedBy: LabelManagedByValue}, + OwnerReferences: []metav1.OwnerReference{ownerReference(dep)}, + }, + Data: map[string]string{"version": "v1.14.0"}, + } + r := newReconciler(t, nil, dep, status) + + reconcileOnce(t, r) + if _, err := getJob(r); err != nil { + t.Fatalf("expected legacy version-only status to be migrated to the new identity: %v", err) + } +} + +func TestReconcile_StatusOwnedByPreviousDeployment_Reruns(t *testing.T) { + oldDep := newTestDeployment("openfga/openfga:v1.14.0") + oldDep.UID = "old-deployment-uid" + status := newStatus(oldDep) + dep := newTestDeployment("openfga/openfga:v1.14.0") + r := newReconciler(t, nil, dep, status) + + reconcileOnce(t, r) + if _, err := getJob(r); err != nil { + t.Fatalf("expected a new Deployment to rerun the migration: %v", err) + } +} + +func TestReconcile_UnownedStatusCollision_ReturnsError(t *testing.T) { + dep := newTestDeployment("openfga/openfga:v1.14.0") + status := &corev1.ConfigMap{ + ObjectMeta: metav1.ObjectMeta{Name: statusKey.Name, Namespace: statusKey.Namespace}, + Data: map[string]string{"version": "v1.14.0"}, + } + r := newReconciler(t, nil, dep, status) + + if _, err := r.Reconcile(context.Background(), ctrl.Request{NamespacedName: deploymentKey}); err == nil { + t.Fatal("expected an unowned status ConfigMap collision to return an error") + } + if _, err := getJob(r); !apierrors.IsNotFound(err) { + t.Errorf("the collision must prevent migration Job creation, got err=%v", err) + } +} + func TestReconcile_StartedJobWithOutdatedTemplate_Kept(t *testing.T) { for _, tt := range []struct { name string @@ -422,6 +561,73 @@ func TestReconcile_StartedJobWithOutdatedTemplate_Kept(t *testing.T) { } } +func TestReconcile_RunningJobForPreviousVersion_Kept(t *testing.T) { + oldDep := newTestDeployment("openfga/openfga:v1.14.0") + job := newTestJob(oldDep) + job.Status.Active = 1 + job.Status.Ready = ptr.To(int32(1)) + r := newReconciler(t, nil, newTestDeployment("openfga/openfga:v1.15.0"), job) + + if result := reconcileOnce(t, r); result.RequeueAfter != 10*time.Second { + t.Errorf("expected the running migration to be polled, got %v", result.RequeueAfter) + } + kept, err := getJob(r) + if err != nil { + t.Fatalf("expected the running migration to be kept: %v", err) + } + if kept.Annotations[AnnotationDesiredVersion] != "v1.14.0" { + t.Errorf("expected the v1.14.0 migration to finish, got %q", kept.Annotations[AnnotationDesiredVersion]) + } +} + +func TestReconcile_UnownedJobCollision_ReturnsError(t *testing.T) { + dep := newTestDeployment("openfga/openfga:v1.14.0") + external := &batchv1.Job{ObjectMeta: metav1.ObjectMeta{Name: jobKey.Name, Namespace: jobKey.Namespace}} + r := newReconciler(t, nil, dep, external) + + if _, err := r.Reconcile(context.Background(), ctrl.Request{NamespacedName: deploymentKey}); err == nil { + t.Fatal("expected an unowned Job collision to return an error") + } + if _, err := getJob(r); err != nil { + t.Errorf("the unowned Job must not be deleted: %v", err) + } +} + +func TestReconcile_JobOwnedByPreviousDeployment_ReplacedWithPreconditions(t *testing.T) { + oldDep := newTestDeployment("openfga/openfga:v1.14.0") + oldDep.UID = "old-deployment-uid" + job := newTestJob(oldDep) + job.UID = "old-job-uid" + job.ResourceVersion = "7" + dep := newTestDeployment("openfga/openfga:v1.14.0") + var checkedPreconditions bool + r := newReconciler(t, &interceptor.Funcs{ + Delete: func(ctx context.Context, c client.WithWatch, obj client.Object, opts ...client.DeleteOption) error { + applied := (&client.DeleteOptions{}).ApplyOptions(opts) + if applied.Preconditions == nil || + applied.Preconditions.UID == nil || + *applied.Preconditions.UID != job.UID || + applied.Preconditions.ResourceVersion == nil || + *applied.Preconditions.ResourceVersion != job.ResourceVersion { + return fmt.Errorf("missing delete preconditions: %+v", applied.Preconditions) + } + if applied.PropagationPolicy == nil || *applied.PropagationPolicy != metav1.DeletePropagationForeground { + return fmt.Errorf("expected foreground deletion, got %v", applied.PropagationPolicy) + } + checkedPreconditions = true + return c.Delete(ctx, obj, opts...) + }, + }, dep, job) + + reconcileOnce(t, r) + if !checkedPreconditions { + t.Fatal("expected Job deletion to include UID and resourceVersion preconditions") + } + if _, err := getJob(r); !apierrors.IsNotFound(err) { + t.Errorf("expected the stale operator Job to be deleted, got err=%v", err) + } +} + // The legacy chart's Helm hook Job has the same name and carries the chart's // app.kubernetes.io/version label, which is the chart appVersion rather than // the image it ran. It must never be taken as proof of a migration. @@ -507,6 +713,80 @@ func TestReconcile_ContainerNotFound_ReturnsError(t *testing.T) { } } +func TestBuildMigrationJob_UsesMigrationPodConfiguration(t *testing.T) { + dep := newTestDeployment("openfga/openfga:v1.14.0") + dep.Spec.Template.Spec.Volumes = []corev1.Volume{{ + Name: "shared", + VolumeSource: corev1.VolumeSource{ + EmptyDir: &corev1.EmptyDirVolumeSource{}, + }, + }} + dep.Spec.Template.Spec.Containers[0].VolumeMounts = []corev1.VolumeMount{{Name: "shared", MountPath: "/shared"}} + dep.Annotations[AnnotationMigrationInitContainers] = `[{"name":"prepare-proxy","image":"busybox:1.36"}]` + dep.Annotations[AnnotationMigrationSidecars] = `[{"name":"database-proxy","image":"proxy:v2"}]` + dep.Annotations[AnnotationMigrationVolumes] = `[{"name":"credentials","secret":{"secretName":"database-proxy"}}]` + dep.Annotations[AnnotationMigrationVolumeMounts] = `[{"name":"credentials","mountPath":"/credentials","readOnly":true}]` + dep.Annotations[AnnotationMigrationResources] = `{"requests":{"cpu":"100m"}}` + dep.Annotations[AnnotationMigrationTimeout] = "2m" + dep.Annotations[AnnotationMigrationNonce] = "secret-rotation-2" + dep.Annotations[AnnotationMigrationAnnotations] = `{"admission.example.com/inject":"enabled","openfga.dev/migration-nonce":"ignored"}` + dep.Annotations[AnnotationMigrationLabels] = `{"app.kubernetes.io/component":"overridden","network-policy.example.com/database":"allowed"}` + + job := newTestJob(dep) + spec := job.Spec.Template.Spec + if len(spec.InitContainers) != 1 || spec.InitContainers[0].Name != "prepare-proxy" { + t.Errorf("unexpected init containers: %+v", spec.InitContainers) + } + if len(spec.Containers) != 2 || spec.Containers[1].Name != "database-proxy" { + t.Errorf("unexpected containers: %+v", spec.Containers) + } + if len(spec.Volumes) != 2 || spec.Volumes[1].Name != "credentials" { + t.Errorf("unexpected volumes: %+v", spec.Volumes) + } + migrate := spec.Containers[0] + if len(migrate.VolumeMounts) != 2 || migrate.VolumeMounts[1].MountPath != "/credentials" { + t.Errorf("unexpected migration volume mounts: %+v", migrate.VolumeMounts) + } + if got := migrate.Resources.Requests[corev1.ResourceCPU]; got.Cmp(resource.MustParse("100m")) != 0 { + t.Errorf("expected 100m CPU request, got %s", got.String()) + } + if !hasEnvVar(migrate.Env, "OPENFGA_TIMEOUT") { + t.Errorf("expected OPENFGA_TIMEOUT in %+v", migrate.Env) + } + if got := job.Spec.Template.Annotations[AnnotationMigrationNonce]; got != "secret-rotation-2" { + t.Errorf("expected migration nonce on the Job pod template, got %q", got) + } + if got := job.Annotations["admission.example.com/inject"]; got != "enabled" { + t.Errorf("expected custom Job annotation, got %q", got) + } + if got := job.Spec.Template.Annotations["admission.example.com/inject"]; got != "enabled" { + t.Errorf("expected custom pod annotation, got %q", got) + } + if got := job.Labels[LabelComponent]; got != "migration" { + t.Errorf("operator Job label must take precedence, got %q", got) + } + if got := job.Spec.Template.Labels[LabelComponent]; got != "migration" { + t.Errorf("operator pod label must take precedence, got %q", got) + } + if got := job.Spec.Template.Labels["network-policy.example.com/database"]; got != "allowed" { + t.Errorf("expected custom pod label, got %q", got) + } +} + +func TestReconcile_InvalidMigrationPodConfiguration_ReturnsError(t *testing.T) { + for _, value := range []string{`not-json`, `[{"name":"proxy","image":"proxy:v2","restartPolcy":"Always"}]`, `[] {}`} { + t.Run(value, func(t *testing.T) { + dep := newTestDeployment("openfga/openfga:v1.14.0") + dep.Annotations[AnnotationMigrationSidecars] = value + r := newReconciler(t, nil, dep) + + if _, err := r.Reconcile(context.Background(), ctrl.Request{NamespacedName: deploymentKey}); err == nil { + t.Fatal("expected invalid migration sidecars to return an error") + } + }) + } +} + func TestExtractImageTag(t *testing.T) { tests := []struct { image string From 9c4a1e152c0660ac69771f49646da6bb6836946b Mon Sep 17 00:00:00 2001 From: SoulPancake <angbpy@gmail.com> Date: Wed, 23 Sep 2026 16:42:05 +0530 Subject: [PATCH 63/70] operator: build the Job from the whole pod spec, key migrations on the datastore too Review findings on the migration Job: - Sidecars and init containers: the Job now runs the Deployment's other containers as native sidecars (restartPolicy: Always, so they stop when migrate exits) and copies its init containers, so a database proxy on localhost works. Needs Kubernetes 1.29 when sidecars are present. - migrate.labels and migrate.annotations are forwarded through the openfga.dev/migration-labels and -annotations Deployment annotations, with helm.sh/* keys dropped, so Istio/Linkerd injection can be disabled for the migration pod. The operator's identity labels and openfga.dev/* annotations cannot be overridden. - The migration identity is the image tag plus an openfga.dev/migration-trigger annotation the chart derives from the datastore settings and the new migration.trigger value, so pointing the release at another database with the same image runs the migration, and a rotated Secret can be handled declaratively instead of by deleting the status ConfigMap. - Only Jobs the operator created or the legacy hook Job are replaced; a same-name Job from elsewhere is left alone and reported with a MigrationJobConflict event. A status ConfigMap not owned by the Deployment is not trusted. Deletion uses foreground propagation with UID and resourceVersion preconditions so two migrations never overlap. - ttlSecondsAfterFinished is applied after success only. Set at creation, a TTL below the retry delay removed failed Jobs before the retry, and a TTL of 0 could remove a completed Job before its outcome was recorded. - Events on the Deployment for started, succeeded, failed and conflicting migrations. The release guard moves to .github/scripts/check-operator-release.sh and also requires the openfga chart version and its operator dependency to move. --- .github/scripts/check-operator-release.sh | 45 ++++ .github/workflows/operator.yml | 24 +- charts/openfga-operator/templates/role.yaml | 2 +- charts/openfga/README.md | 2 +- charts/openfga/templates/_helpers.tpl | 23 ++ charts/openfga/templates/deployment.yaml | 8 + charts/openfga/tests/operator_mode_test.yaml | 52 ++++ charts/openfga/values.schema.json | 7 +- charts/openfga/values.yaml | 9 +- docs/adr/002-operator-managed-migrations.md | 6 +- operator/README.md | 25 +- operator/cmd/main.go | 1 + operator/internal/controller/helpers.go | 144 +++++++++-- .../controller/migration_controller.go | 106 ++++++-- .../controller/migration_controller_test.go | 243 +++++++++++++++++- 15 files changed, 606 insertions(+), 91 deletions(-) create mode 100755 .github/scripts/check-operator-release.sh diff --git a/.github/scripts/check-operator-release.sh b/.github/scripts/check-operator-release.sh new file mode 100755 index 00000000..b1cd64d1 --- /dev/null +++ b/.github/scripts/check-operator-release.sh @@ -0,0 +1,45 @@ +#!/usr/bin/env bash +# Fails when the operator image changed but the chart versions that publish it +# did not. Usage: check-operator-release.sh <base-ref>, e.g. origin/main. +# +# CI publishes ghcr.io/openfga/openfga-operator:<appVersion> once and never +# overwrites it, and chart-releaser skips chart versions that already exist. A +# change to the image is therefore only released when, in the same PR: +# 1. charts/openfga-operator/Chart.yaml bumps appVersion and version +# 2. charts/openfga/Chart.yaml bumps version and pins the new operator chart +# Chart.lock has to match too; `helm dependency build` in CI fails if it does not. +set -euo pipefail + +base=$1 +image_inputs=(operator/cmd operator/internal operator/go.mod operator/go.sum operator/Dockerfile) +if git diff --quiet "$base...HEAD" -- "${image_inputs[@]}"; then + echo "operator image inputs unchanged" + exit 0 +fi + +field() { grep "^$1:" "$2" | awk '{print $2}' | tr -d '"'; } +dependency_version() { awk '/name: openfga-operator/{f=1} f && /version:/{print $2; exit}' "$1" | tr -d '"'; } +at_base() { git show "$base:$1" 2>/dev/null; } +fail() { echo "::error file=$1::$2"; exit 1; } + +operator_chart=charts/openfga-operator/Chart.yaml +parent_chart=charts/openfga/Chart.yaml + +if at_base "$operator_chart" > /tmp/base-operator-chart.yaml; then + for f in appVersion version; do + if [[ "$(field $f /tmp/base-operator-chart.yaml)" == "$(field $f $operator_chart)" ]]; then + fail "$operator_chart" "operator image inputs changed but $f is still $(field $f $operator_chart); bump appVersion and version so the change is published" + fi + done +else + echo "operator chart is new in this PR" +fi + +at_base "$parent_chart" > /tmp/base-parent-chart.yaml +if [[ "$(field version /tmp/base-parent-chart.yaml)" == "$(field version $parent_chart)" ]]; then + fail "$parent_chart" "operator image inputs changed but the openfga chart version is still $(field version $parent_chart); bump it so a chart with the new operator is released" +fi +if [[ "$(dependency_version $parent_chart)" != "$(field version $operator_chart)" ]]; then + fail "$parent_chart" "openfga chart pins openfga-operator $(dependency_version $parent_chart) but the operator chart is $(field version $operator_chart); update the dependency and run helm dependency update charts/openfga" +fi +echo "operator release versions are consistent" diff --git a/.github/workflows/operator.yml b/.github/workflows/operator.yml index bc1f79d9..fd1a796c 100644 --- a/.github/workflows/operator.yml +++ b/.github/workflows/operator.yml @@ -36,29 +36,7 @@ jobs: - name: Require a version bump when the operator image changes if: github.event_name == 'pull_request' - env: - BASE: ${{ github.base_ref }} - run: | - set -euo pipefail - # CI publishes ghcr.io/openfga/openfga-operator:<appVersion> once and never - # overwrites it, so a change to the image is only released when the - # operator chart's appVersion and version are bumped. - image_inputs=(operator/cmd operator/internal operator/go.mod operator/go.sum operator/Dockerfile) - if git diff --quiet "origin/$BASE...HEAD" -- "${image_inputs[@]}"; then - echo "operator image inputs unchanged" - exit 0 - fi - field() { grep "^$1:" "$2" | awk '{print $2}' | tr -d '"'; } - if ! git show "origin/$BASE:charts/openfga-operator/Chart.yaml" > /tmp/base-chart.yaml 2>/dev/null; then - echo "operator chart is new in this PR" - exit 0 - fi - for f in appVersion version; do - if [[ "$(field $f /tmp/base-chart.yaml)" == "$(field $f charts/openfga-operator/Chart.yaml)" ]]; then - echo "::error file=charts/openfga-operator/Chart.yaml::operator image inputs changed but $f is still $(field $f charts/openfga-operator/Chart.yaml); bump appVersion and version so the change is published" - exit 1 - fi - done + run: .github/scripts/check-operator-release.sh "origin/${{ github.base_ref }}" - name: Set up Go uses: actions/setup-go@40f1582b2485089dde7abd97c1529aa768e1baff # v5.6.0 diff --git a/charts/openfga-operator/templates/role.yaml b/charts/openfga-operator/templates/role.yaml index e9e851f4..ef4bc5dc 100644 --- a/charts/openfga-operator/templates/role.yaml +++ b/charts/openfga-operator/templates/role.yaml @@ -14,7 +14,7 @@ rules: verbs: ["patch"] - apiGroups: ["batch"] resources: ["jobs"] - verbs: ["get", "list", "watch", "create", "delete"] + verbs: ["get", "list", "watch", "create", "patch", "delete"] - apiGroups: [""] resources: ["configmaps"] verbs: ["get", "list", "watch", "create", "update"] diff --git a/charts/openfga/README.md b/charts/openfga/README.md index f9ae8971..f4d3b8e6 100644 --- a/charts/openfga/README.md +++ b/charts/openfga/README.md @@ -164,7 +164,7 @@ datastore: uriSecret: my-postgres-secret ``` -The operator only runs migrations; replicas, autoscaling and the pod template stay under the chart's control. It records the migrated version in the `<release>-migration-status` ConfigMap and sets a `MigrationFailed` condition on the Deployment if a migration fails. The migration Job runs as a dedicated `<release>-migration` service account (`migration.serviceAccount`), which can carry cloud IAM annotations for DDL permissions. Migrations only run when the image tag changes, so pin `image.tag` to a release rather than a floating tag. See the [operator README](../../operator/README.md) for how it works and its limitations. +The operator only runs migrations; replicas, autoscaling and the pod template stay under the chart's control. It records the migrated version in the `<release>-migration-status` ConfigMap and sets a `MigrationFailed` condition on the Deployment if a migration fails. The migration Job is built from the OpenFGA pod spec, so `sidecars` such as a database proxy and `extraInitContainers` run alongside it, and `migrate.labels`/`migrate.annotations` (e.g. `sidecar.istio.io/inject: "false"`) are applied to it. It runs as a dedicated `<release>-migration` service account (`migration.serviceAccount`), which can carry cloud IAM annotations for DDL permissions. Migrations run when the image tag or the datastore settings change, so pin `image.tag` to a release rather than a floating tag; set `migration.trigger` to any new value to run one on demand. See the [operator README](../../operator/README.md) for how it works and its limitations. ## Uninstalling the Chart diff --git a/charts/openfga/templates/_helpers.tpl b/charts/openfga/templates/_helpers.tpl index 46b8876a..baea73c1 100644 --- a/charts/openfga/templates/_helpers.tpl +++ b/charts/openfga/templates/_helpers.tpl @@ -85,6 +85,29 @@ Create the name of the migration service account to use (operator mode only) {{- end }} {{- end }} +{{/* +Identity of the datastore the operator migrates. The operator runs the migration +again whenever this changes, e.g. when the chart points at a different database +with the same OpenFGA image. migration.trigger is any user-chosen string that +forces another run, for cases the chart cannot see such as a Secret rotation. +*/}} +{{- define "openfga.migrationTrigger" -}} +{{- $ds := .Values.datastore -}} +{{- dict "engine" $ds.engine "uri" $ds.uri "uriSecret" $ds.uriSecret "username" $ds.username "password" $ds.password "existingSecret" $ds.existingSecret "secretKeys" $ds.secretKeys "trigger" .Values.migration.trigger | toJson | sha256sum | trunc 16 -}} +{{- end -}} + +{{/* +migrate.annotations without Helm hook keys, as JSON for the operator to put on +the migration Job and its pod +*/}} +{{- define "openfga.migrationAnnotations" -}} +{{- $out := dict -}} +{{- range $k, $v := .Values.migrate.annotations -}} +{{- if not (hasPrefix "helm.sh/" $k) -}}{{- $_ := set $out $k (toString $v) -}}{{- end -}} +{{- end -}} +{{- toJson $out -}} +{{- end -}} + {{/* Return true if the openfga-operator runs the database migrations for this release */}} diff --git a/charts/openfga/templates/deployment.yaml b/charts/openfga/templates/deployment.yaml index 6c81a287..61296700 100644 --- a/charts/openfga/templates/deployment.yaml +++ b/charts/openfga/templates/deployment.yaml @@ -10,9 +10,17 @@ metadata: {{- if $operatorMigrations }} openfga.dev/migration-enabled: "true" openfga.dev/container-name: "{{ .Chart.Name }}" + openfga.dev/migration-trigger: {{ include "openfga.migrationTrigger" . | quote }} {{- if or .Values.migration.serviceAccount.create .Values.migration.serviceAccount.name }} openfga.dev/migration-service-account: '{{ include "openfga.migrationServiceAccountName" . }}' {{- end }} + {{- with .Values.migrate.labels }} + openfga.dev/migration-labels: {{ toJson . | quote }} + {{- end }} + {{- $migrationAnnotations := include "openfga.migrationAnnotations" . }} + {{- if ne $migrationAnnotations "{}" }} + openfga.dev/migration-annotations: {{ $migrationAnnotations | quote }} + {{- end }} {{- end }} {{- with .Values.annotations }} {{- toYaml . | nindent 4 }} diff --git a/charts/openfga/tests/operator_mode_test.yaml b/charts/openfga/tests/operator_mode_test.yaml index 16dbc197..d5868f81 100644 --- a/charts/openfga/tests/operator_mode_test.yaml +++ b/charts/openfga/tests/operator_mode_test.yaml @@ -38,6 +38,58 @@ tests: - isNull: path: metadata.annotations + - it: should derive the migration trigger from the datastore settings + set: + openfga-operator.enabled: true + datastore.engine: postgres + datastore.uri: postgres://a/openfga + asserts: + - matchRegex: + path: metadata.annotations["openfga.dev/migration-trigger"] + pattern: ^[0-9a-f]{16}$ + + - it: should change the migration trigger when the datastore or migration.trigger changes + set: + openfga-operator.enabled: true + datastore.engine: postgres + datastore.uri: postgres://b/openfga + migration.trigger: "2" + asserts: + - matchRegex: + path: metadata.annotations["openfga.dev/migration-trigger"] + pattern: ^[0-9a-f]{16}$ + - notEqual: + path: metadata.annotations["openfga.dev/migration-trigger"] + # value for postgres://a/openfga with no trigger, from the test above + value: 852c9a016f0ffb87 + + - it: should forward migrate labels and non-hook annotations to the operator + set: + openfga-operator.enabled: true + datastore.engine: postgres + migrate.labels: + team: auth + migrate.annotations: + helm.sh/hook: post-install + sidecar.istio.io/inject: "false" + asserts: + - equal: + path: metadata.annotations["openfga.dev/migration-labels"] + value: '{"team":"auth"}' + - equal: + path: metadata.annotations["openfga.dev/migration-annotations"] + value: '{"sidecar.istio.io/inject":"false"}' + + - it: should not emit migration metadata annotations for hook-only annotations + set: + openfga-operator.enabled: true + datastore.engine: postgres + asserts: + - isNull: + path: metadata.annotations["openfga.dev/migration-labels"] + - isNull: + path: metadata.annotations["openfga.dev/migration-annotations"] + - it: should use custom migration service account name when set set: openfga-operator.enabled: true diff --git a/charts/openfga/values.schema.json b/charts/openfga/values.schema.json index bc56e79a..f00aa4bb 100644 --- a/charts/openfga/values.schema.json +++ b/charts/openfga/values.schema.json @@ -1306,8 +1306,13 @@ }, "migration": { "type": "object", - "description": "Service account for the migration Jobs the operator creates. Only used when openfga-operator.enabled is true.", + "description": "Settings for the migration Jobs the operator creates. Only used when openfga-operator.enabled is true.", "properties": { + "trigger": { + "type": "string", + "description": "Change to any new value to run the migration again without changing the image or datastore settings", + "default": "" + }, "serviceAccount": { "type": "object", "properties": { diff --git a/charts/openfga/values.yaml b/charts/openfga/values.yaml index acef64f1..ce302c43 100644 --- a/charts/openfga/values.yaml +++ b/charts/openfga/values.yaml @@ -372,10 +372,13 @@ migrate: extraVolumeMounts: [] extraInitContainers: [] sidecars: [] + # -- Added to the migration Job and its pod. With the operator, helm.sh/* keys + # are dropped and the rest is forwarded (e.g. sidecar.istio.io/inject: "false"). annotations: helm.sh/hook: "post-install, post-upgrade, post-rollback, post-delete" helm.sh/hook-weight: "-5" helm.sh/hook-delete-policy: "before-hook-creation" + # -- Added to the migration Job and its pod in both modes. labels: {} timeout: @@ -405,9 +408,13 @@ openfga-operator: # limits: # memory: 128Mi -# -- Service account for the migration Jobs the operator creates. +# -- Settings for the migration Jobs the operator creates. # Only used when openfga-operator.enabled is true. migration: + # -- The operator runs a migration when the OpenFGA image or the datastore + # settings change. Set this to any new value to run it again when neither + # changed, e.g. after rotating a Secret to point at a different database. + trigger: "" serviceAccount: # -- Create a dedicated service account for migration Jobs. # The migration Job inherits env vars (including secretKeyRef) from the OpenFGA container. diff --git a/docs/adr/002-operator-managed-migrations.md b/docs/adr/002-operator-managed-migrations.md index 72e94bc4..1cf203e7 100644 --- a/docs/adr/002-operator-managed-migrations.md +++ b/docs/adr/002-operator-managed-migrations.md @@ -103,7 +103,7 @@ Readiness comes from OpenFGA itself: `IsReady()` reports `NOT_SERVING` while the #### Version tracking via ConfigMap -A ConfigMap (`openfga-migration-status`) records the last successfully migrated version. The operator compares this to the Deployment's image tag to determine if migration is needed. This is: +A ConfigMap (`openfga-migration-status`) records the last successfully migrated identity: the image version and the datastore trigger. The operator compares this to the Deployment to determine if migration is needed. This is: - Simple to inspect (`kubectl get configmap openfga-migration-status -o yaml`) - Survives operator restarts - Can be manually deleted to force re-migration (once the previous migration Job has been cleaned up) @@ -122,7 +122,9 @@ The Job created by the operator has no Helm hook annotations. It is a standard K |---------|----------| | Job fails | Operator sets `MigrationFailed` on the Deployment, keeps the failed Job for 60 seconds so its logs can be read, then replaces it. On a fresh database the pods stay `NotReady`; on an upgrade they keep serving on the previous schema. | | Job pod never starts | A bad secret reference, image pull error or unschedulable pod never fails the Job. Once the Deployment's pod template changes (the fix rolls out), the operator rebuilds a Job whose pod is not running. | -| Image changes while a Job runs | The running Job is left to finish and then replaced by one for the new image. The hook flow deletes the running hook Job instead (`before-hook-creation`), which can abort a concurrent index build and leave it invalid. | +| Image changes while a Job runs | The running Job is left to finish and then replaced by one for the new image. The hook flow deletes the running hook Job instead (`before-hook-creation`), which can abort a concurrent index build and leave it invalid. Replacement uses foreground deletion, so the next Job only starts once the old pods are gone. | +| Datastore repointed, same image | The chart derives a trigger from the datastore settings that is part of the migration identity, so the migration runs against the new database. `migration.trigger` forces a run for changes the chart cannot see. | +| Same-name Job or ConfigMap from elsewhere | Not touched; a `MigrationJobConflict` event is recorded until it is removed. | | Job hangs | No deadline by default, like the Helm hook Job. `activeDeadlineSeconds` can be set, but a migration cut off halfway (an index build, a MySQL table rebuild) starts over on the next attempt. | | Operator crashes | On restart, re-reads the ConfigMap and Job status and resumes. The retry delay is measured from the failed Job's condition, so it survives restarts. | | Database unreachable | Job fails to connect. After exhausting `backoffLimit` the cycle above repeats until the database becomes available. | diff --git a/operator/README.md b/operator/README.md index 24778f3d..9e251544 100644 --- a/operator/README.md +++ b/operator/README.md @@ -7,10 +7,10 @@ This is **Stage 1** of the operator — focused solely on migration orchestratio ## How It Works 1. The operator watches Deployments in its configured namespace, which defaults to the operator pod's namespace, labeled `app.kubernetes.io/part-of: openfga` and `app.kubernetes.io/component: authorization-controller` -2. When a version change is detected (comparing the container image tag to the `{name}-migration-status` ConfigMap), the operator: - - Creates a migration Job running `openfga migrate` +2. When the migration identity changes (the container image tag plus the `openfga.dev/migration-trigger` annotation, compared to the `{name}-migration-status` ConfigMap), the operator: + - Creates a migration Job running `openfga migrate`, built from the Deployment's pod spec: the OpenFGA container's image, env, volumes, resources and scheduling, the other containers as [sidecars](https://kubernetes.io/docs/concepts/workloads/pods/sidecar-containers/) that stop when the migration exits, and the init containers - Waits for the Job to complete - - Updates the ConfigMap with the new version + - Records the identity in the ConfigMap and applies `ttlSecondsAfterFinished` so Kubernetes cleans the Job up 3. On failure, a `MigrationFailed` condition is set on the Deployment. The failed Job is kept for 60 seconds so its logs can be inspected, then replaced with a new one. A running migration is never interrupted. If the image changes again while a Job's pod is running (a rollback, or two upgrades in a row), the operator waits for that Job to finish and then runs the migration for the new image. Aborting a non-transactional step such as Postgres's concurrent index build in migration 006 leaves an invalid index that the next run skips. To abort a migration that is stuck, delete the Job or set `migrationJob.activeDeadlineSeconds`. @@ -22,7 +22,7 @@ The operator never changes the Deployment's replica count or pod template. On a - Go 1.26.8+ - Docker - Helm 3.6+ -- A Kubernetes cluster (Rancher Desktop, kind, etc.) +- A Kubernetes cluster (Rancher Desktop, kind, etc.), 1.29 or newer when the OpenFGA pod has sidecars ## Development @@ -55,8 +55,10 @@ docker build -t openfga/openfga-operator:dev . CI publishes `ghcr.io/openfga/openfga-operator:<appVersion>` on the first push to `main` that carries that appVersion and never overwrites it, and chart-releaser likewise skips chart versions that already exist. A change to the operator image (`cmd/`, `internal/`, `go.mod`, `go.sum`, `Dockerfile`) therefore has to bump, in the same PR: -1. `appVersion` and `version` in `charts/openfga-operator/Chart.yaml` (the operator workflow fails the PR otherwise) -2. the `openfga-operator` dependency version and `version` in `charts/openfga/Chart.yaml`, then `helm dependency update charts/openfga` to refresh `Chart.lock` (`helm dependency build` fails otherwise) +1. `appVersion` and `version` in `charts/openfga-operator/Chart.yaml` +2. the `openfga-operator` dependency version and `version` in `charts/openfga/Chart.yaml`, then `helm dependency update charts/openfga` to refresh `Chart.lock` + +`.github/scripts/check-operator-release.sh origin/main` checks the first two files and runs on every PR; `helm dependency build` fails when `Chart.lock` is stale. ## Local Testing @@ -135,12 +137,15 @@ The operator reads these annotations from the OpenFGA Deployment: |------------|-------------| | `openfga.dev/migration-enabled` | Must be `"true"` for the operator to manage migrations. Deployments without this annotation are ignored. Set by the Helm chart when `openfga-operator.enabled` and `datastore.applyMigrations` are true and the datastore is Postgres or MySQL. | | `openfga.dev/container-name` | The OpenFGA container in the pod spec. Defaults to `openfga`. | +| `openfga.dev/migration-trigger` | Any string that is part of the migration identity alongside the image tag; a change runs the migration again. The chart derives it from the datastore settings and `migration.trigger`. | | `openfga.dev/migration-service-account` | The ServiceAccount to use for migration Jobs. Defaults to the Deployment's SA. | +| `openfga.dev/migration-labels`, `openfga.dev/migration-annotations` | JSON maps added to the migration Job and its pod, e.g. `{"sidecar.istio.io/inject":"false"}`. The chart fills them from `migrate.labels` and `migrate.annotations` (without `helm.sh/*` keys). The operator's own labels and `openfga.dev/*` annotations cannot be overridden. | ## Limitations -- **Migrations key only on the image tag:** The operator compares the container image tag (or digest) to the `{name}-migration-status` ConfigMap. A mutable tag like `latest`, or a tag reused for a new build, is not seen as a change, so the migration is skipped — use immutable tags (e.g. `v1.14.0`) or pin by digest. A migration-needing change that keeps the same image — for example repointing `datastore.uri` at a different or restored database — also won't trigger a Job; delete the status ConfigMap (and the `{name}-migrate` Job, if it still exists) to run the migration again. -- **Legacy migration values:** `migrate.*` (extra volumes and mounts, init containers, sidecars, annotations, labels, timeout) and `datastore.migrations.resources` only apply to the legacy Helm hook Job. The operator's Job copies the OpenFGA container's image, env, volumes, resources, security context and scheduling instead, so put anything the migration needs (e.g. CA bundles) in the top-level `extraVolumes`, `extraVolumeMounts` and `extraEnvVars`. -- **Single-container migration Job:** The Job runs one container (`openfga migrate`) and injects no sidecars or extra init containers, so databases reached through a sidecar proxy (Cloud SQL Auth Proxy, AlloyDB) aren't supported for operator-managed migrations. A sidecar injected into every pod in the namespace that doesn't exit on its own (e.g. an Istio sidecar) keeps the Job pod running and stops the Job from completing. -- **Job pod labels:** The migration pod is labelled `app.kubernetes.io/part-of: openfga` and `app.kubernetes.io/component: migration`, not with the OpenFGA Deployment's `app.kubernetes.io/name`/`instance` labels (which would make it a Service endpoint). A NetworkPolicy that allows database egress only for the OpenFGA pods' labels needs a rule for the migration pod too. +- **Migrations key on the image tag and the trigger:** A mutable tag like `latest`, or a tag reused for a new build, is not seen as a change, so use immutable tags (e.g. `v1.14.0`) or pin by digest. The chart's trigger covers changes to the datastore settings it renders, but not a Secret whose contents change under the same name; set `migration.trigger` to a new value in that case. +- **Legacy migration values:** `migrate.extraVolumes`, `migrate.extraVolumeMounts`, `migrate.extraInitContainers`, `migrate.sidecars`, `migrate.timeout` and `datastore.migrations.resources` only apply to the legacy Helm hook Job. The operator's Job is built from the OpenFGA pod spec, so put what the migration needs in the top-level `extraVolumes`, `extraVolumeMounts`, `extraInitContainers`, `sidecars` and `extraEnvVars`. `migrate.labels` and `migrate.annotations` apply in both modes. +- **Injected sidecars:** Containers injected by a webhook (Istio, Linkerd) are not part of the Deployment's pod spec and are not converted to native sidecars, so one that does not exit keeps the Job pod running. Disable injection for the migration pod with `migrate.annotations`, e.g. `sidecar.istio.io/inject: "false"`. +- **Job pod labels:** The migration pod is labelled `app.kubernetes.io/part-of: openfga` and `app.kubernetes.io/component: migration` plus `migrate.labels`, not with the OpenFGA Deployment's `app.kubernetes.io/name`/`instance` labels (which would make it a Service endpoint). A NetworkPolicy that allows database egress only for the OpenFGA pods' labels needs a rule for the migration pod too. +- **Same-name resources:** The operator only replaces a `{name}-migrate` Job it created itself or the chart's legacy hook Job, and only trusts a `{name}-migration-status` ConfigMap owned by the Deployment. Anything else with those names blocks the migration with a `MigrationJobConflict` event until it is removed. - **One namespace per operator:** The operator reconciles every opted-in OpenFGA Deployment in its watch namespace. Operators installed by several releases in one namespace share a leader election lease, so only one of them is active at a time. diff --git a/operator/cmd/main.go b/operator/cmd/main.go index 783196db..40f96fa1 100644 --- a/operator/cmd/main.go +++ b/operator/cmd/main.go @@ -97,6 +97,7 @@ func main() { reconciler := &controller.MigrationReconciler{ Client: mgr.GetClient(), + Recorder: mgr.GetEventRecorderFor("openfga-operator"), BackoffLimit: int32(backoffLimit), ActiveDeadlineSeconds: int64(activeDeadline), TTLSecondsAfterFinished: int32(ttlAfterFinished), diff --git a/operator/internal/controller/helpers.go b/operator/internal/controller/helpers.go index a82f1799..0f6ef578 100644 --- a/operator/internal/controller/helpers.go +++ b/operator/internal/controller/helpers.go @@ -5,6 +5,7 @@ import ( "crypto/sha256" "encoding/json" "fmt" + "maps" "strings" "time" @@ -33,9 +34,15 @@ const ( AnnotationMigrationEnabled = "openfga.dev/migration-enabled" AnnotationContainerName = "openfga.dev/container-name" AnnotationMigrationServiceAccount = "openfga.dev/migration-service-account" + // AnnotationMigrationTrigger is any string that, together with the image + // version, identifies a migration. Changing it runs the migration again. + AnnotationMigrationTrigger = "openfga.dev/migration-trigger" + // AnnotationMigrationLabels and AnnotationMigrationAnnotations hold JSON + // maps of metadata to add to the migration Job and its pod. + AnnotationMigrationLabels = "openfga.dev/migration-labels" + AnnotationMigrationAnnotations = "openfga.dev/migration-annotations" - // Annotations set on migration Jobs: the version the Job migrates to, and a - // hash of the pod template it was built from. + // Annotations set on migration Jobs. AnnotationDesiredVersion = "openfga.dev/desired-version" AnnotationPodTemplateHash = "openfga.dev/pod-template-hash" @@ -46,6 +53,37 @@ const ( DefaultTTLSecondsAfterFinished int32 = 300 ) +// migrationIdentity is what the operator compares to decide whether a +// migration has to run: the OpenFGA image version plus the trigger the chart +// derives from the datastore configuration. +type migrationIdentity struct { + Version string + Trigger string +} + +func desiredIdentity(deployment *appsv1.Deployment, container *corev1.Container) migrationIdentity { + return migrationIdentity{ + Version: extractImageTag(container.Image), + Trigger: deployment.Annotations[AnnotationMigrationTrigger], + } +} + +func jobIdentity(job *batchv1.Job) migrationIdentity { + return migrationIdentity{ + Version: job.Annotations[AnnotationDesiredVersion], + Trigger: job.Annotations[AnnotationMigrationTrigger], + } +} + +// recordedIdentity returns the identity stored in the status ConfigMap, or the +// zero value when the ConfigMap is missing or not owned by this Deployment. +func recordedIdentity(cm *corev1.ConfigMap, deployment *appsv1.Deployment) migrationIdentity { + if !metav1.IsControlledBy(cm, deployment) { + return migrationIdentity{} + } + return migrationIdentity{Version: cm.Data["version"], Trigger: cm.Data["trigger"]} +} + // extractImageTag returns the tag portion of a container image reference. // For "openfga/openfga:v1.14.0" it returns "v1.14.0". // For "openfga/openfga@sha256:abc..." it returns the digest. @@ -89,6 +127,20 @@ func findOpenFGAContainer(deployment *appsv1.Deployment) (*corev1.Container, err return nil, fmt.Errorf("container %q not found in deployment %s/%s", targetName, deployment.Namespace, deployment.Name) } +// metadataFromAnnotation decodes a JSON map of labels or annotations for the +// migration Job from a Deployment annotation. +func metadataFromAnnotation(deployment *appsv1.Deployment, key string) (map[string]string, error) { + raw := deployment.Annotations[key] + if raw == "" { + return nil, nil + } + var m map[string]string + if err := json.Unmarshal([]byte(raw), &m); err != nil { + return nil, fmt.Errorf("parsing annotation %s of deployment %s/%s: %w", key, deployment.Namespace, deployment.Name, err) + } + return m, nil +} + // ownerReference makes the Deployment the controller of a migration Job or // status ConfigMap so both are garbage collected with it. BlockOwnerDeletion // is left unset: it needs update on deployments/finalizers, which the @@ -103,42 +155,82 @@ func ownerReference(deployment *appsv1.Deployment) metav1.OwnerReference { } } +// replaceable reports whether the operator may delete an existing Job that +// has the migration Job's name: one it created for this or an earlier +// Deployment, or the chart's legacy Helm hook Job. Anything else belongs to +// someone else. +func replaceable(job *batchv1.Job, deployment *appsv1.Deployment) bool { + if metav1.IsControlledBy(job, deployment) || job.Labels[LabelManagedBy] == LabelManagedByValue { + return true + } + _, hook := job.Annotations["helm.sh/hook"] + return hook +} + // buildMigrationJob constructs a Job that runs "openfga migrate" with the -// OpenFGA container's image, environment, volumes and scheduling. -func (r *MigrationReconciler) buildMigrationJob(deployment *appsv1.Deployment, container *corev1.Container, version string) *batchv1.Job { +// OpenFGA container's image, environment, volumes and scheduling. The +// Deployment's other containers run as sidecars so a database proxy is +// available to the migration and stops when it exits. +func (r *MigrationReconciler) buildMigrationJob(deployment *appsv1.Deployment, container *corev1.Container, desired migrationIdentity) (*batchv1.Job, error) { + userLabels, err := metadataFromAnnotation(deployment, AnnotationMigrationLabels) + if err != nil { + return nil, err + } + userAnnotations, err := metadataFromAnnotation(deployment, AnnotationMigrationAnnotations) + if err != nil { + return nil, err + } + podSpec := deployment.Spec.Template.Spec serviceAccount := deployment.Annotations[AnnotationMigrationServiceAccount] if serviceAccount == "" { serviceAccount = podSpec.ServiceAccountName } + // Sidecars first so init containers can reach the database through them. + var initContainers []corev1.Container + for i := range podSpec.Containers { + if podSpec.Containers[i].Name == container.Name { + continue + } + sidecar := *podSpec.Containers[i].DeepCopy() + sidecar.RestartPolicy = ptr.To(corev1.ContainerRestartPolicyAlways) + initContainers = append(initContainers, sidecar) + } + initContainers = append(initContainers, podSpec.InitContainers...) + + labels := map[string]string{ + LabelPartOf: LabelPartOfValue, + LabelComponent: "migration", + LabelManagedBy: LabelManagedByValue, + } + podLabels := maps.Clone(userLabels) + if podLabels == nil { + podLabels = map[string]string{} + } + maps.Copy(podLabels, labels) + job := &batchv1.Job{ ObjectMeta: metav1.ObjectMeta{ - Name: migrationJobName(deployment.Name), - Namespace: deployment.Namespace, - Labels: map[string]string{ - LabelPartOf: LabelPartOfValue, - LabelComponent: "migration", - LabelManagedBy: LabelManagedByValue, - }, - Annotations: map[string]string{AnnotationDesiredVersion: version}, + Name: migrationJobName(deployment.Name), + Namespace: deployment.Namespace, + Labels: podLabels, + Annotations: maps.Clone(userAnnotations), OwnerReferences: []metav1.OwnerReference{ownerReference(deployment)}, }, Spec: batchv1.JobSpec{ - BackoffLimit: ptr.To(r.BackoffLimit), - TTLSecondsAfterFinished: ptr.To(r.TTLSecondsAfterFinished), + BackoffLimit: ptr.To(r.BackoffLimit), Template: corev1.PodTemplateSpec{ ObjectMeta: metav1.ObjectMeta{ - Labels: map[string]string{ - LabelPartOf: LabelPartOfValue, - LabelComponent: "migration", - }, + Labels: podLabels, + Annotations: maps.Clone(userAnnotations), }, Spec: corev1.PodSpec{ ServiceAccountName: serviceAccount, RestartPolicy: corev1.RestartPolicyNever, ImagePullSecrets: podSpec.ImagePullSecrets, SecurityContext: podSpec.SecurityContext, + InitContainers: initContainers, Containers: []corev1.Container{{ Name: "migrate-database", Image: container.Image, @@ -161,8 +253,13 @@ func (r *MigrationReconciler) buildMigrationJob(deployment *appsv1.Deployment, c if r.ActiveDeadlineSeconds > 0 { job.Spec.ActiveDeadlineSeconds = ptr.To(r.ActiveDeadlineSeconds) } + if job.Annotations == nil { + job.Annotations = map[string]string{} + } + job.Annotations[AnnotationDesiredVersion] = desired.Version + job.Annotations[AnnotationMigrationTrigger] = desired.Trigger job.Annotations[AnnotationPodTemplateHash] = podTemplateHash(&job.Spec.Template) - return job + return job, nil } func podTemplateHash(template *corev1.PodTemplateSpec) string { @@ -173,8 +270,9 @@ func podTemplateHash(template *corev1.PodTemplateSpec) string { return fmt.Sprintf("%x", sha256.Sum256(b))[:16] } -// updateMigrationStatus records the migrated version in the status ConfigMap. -func updateMigrationStatus(ctx context.Context, c client.Client, deployment *appsv1.Deployment, version, jobName string) error { +// updateMigrationStatus records the migrated identity in the status ConfigMap, +// creating it or taking it over from a previous Deployment of the same name. +func updateMigrationStatus(ctx context.Context, c client.Client, deployment *appsv1.Deployment, identity migrationIdentity, jobName string) error { cm := &corev1.ConfigMap{ObjectMeta: metav1.ObjectMeta{ Name: migrationConfigMapName(deployment.Name), Namespace: deployment.Namespace, @@ -185,10 +283,10 @@ func updateMigrationStatus(ctx context.Context, c client.Client, deployment *app LabelComponent: "migration", LabelManagedBy: LabelManagedByValue, } - // Reset on every write in case the Deployment was recreated with a new UID. cm.OwnerReferences = []metav1.OwnerReference{ownerReference(deployment)} cm.Data = map[string]string{ - "version": version, + "version": identity.Version, + "trigger": identity.Trigger, "migratedAt": time.Now().UTC().Format(time.RFC3339), "jobName": jobName, } diff --git a/operator/internal/controller/migration_controller.go b/operator/internal/controller/migration_controller.go index ba047cb9..15631ea7 100644 --- a/operator/internal/controller/migration_controller.go +++ b/operator/internal/controller/migration_controller.go @@ -11,6 +11,7 @@ import ( apierrors "k8s.io/apimachinery/pkg/api/errors" metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" "k8s.io/apimachinery/pkg/types" + "k8s.io/client-go/tools/record" "k8s.io/utils/ptr" ctrl "sigs.k8s.io/controller-runtime" "sigs.k8s.io/controller-runtime/pkg/builder" @@ -26,12 +27,13 @@ const retryDelay = 60 * time.Second // migration Job whenever the OpenFGA image version changes. type MigrationReconciler struct { client.Client + Recorder record.EventRecorder // BackoffLimit for migration Jobs. BackoffLimit int32 // ActiveDeadlineSeconds for migration Jobs; 0 means no deadline. ActiveDeadlineSeconds int64 - // TTLSecondsAfterFinished for migration Jobs. + // TTLSecondsAfterFinished is applied to migration Jobs once they succeed. TTLSecondsAfterFinished int32 } @@ -51,23 +53,39 @@ func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) ( if err != nil { return ctrl.Result{}, err } - desiredVersion := extractImageTag(container.Image) + desired := desiredIdentity(deployment, container) status := &corev1.ConfigMap{} err = r.Get(ctx, types.NamespacedName{Name: migrationConfigMapName(req.Name), Namespace: req.Namespace}, status) if err != nil && !apierrors.IsNotFound(err) { return ctrl.Result{}, fmt.Errorf("getting migration status: %w", err) } - currentVersion := status.Data["version"] - if currentVersion == desiredVersion { + current := recordedIdentity(status, deployment) + + job := &batchv1.Job{} + err = r.Get(ctx, types.NamespacedName{Name: migrationJobName(req.Name), Namespace: req.Namespace}, job) + if err != nil && !apierrors.IsNotFound(err) { + return ctrl.Result{}, fmt.Errorf("getting migration job: %w", err) + } + if err != nil { + job = nil + } + + if current == desired { + if job != nil && metav1.IsControlledBy(job, deployment) && isJobConditionTrue(job, batchv1.JobComplete) { + if err := r.applyTTL(ctx, job); err != nil { + return ctrl.Result{}, err + } + } _, err := r.patchCondition(ctx, deployment, clearMigrationFailedCondition) return ctrl.Result{}, err } - job := &batchv1.Job{} - err = r.Get(ctx, types.NamespacedName{Name: migrationJobName(req.Name), Namespace: req.Namespace}, job) - if apierrors.IsNotFound(err) { - job = r.buildMigrationJob(deployment, container, desiredVersion) + if job == nil { + job, err = r.buildMigrationJob(deployment, container, desired) + if err != nil { + return ctrl.Result{}, err + } if err := r.Create(ctx, job); err != nil { if apierrors.IsAlreadyExists(err) { // The cache has not caught up with a Job created by an earlier reconcile. @@ -75,27 +93,36 @@ func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) ( } return ctrl.Result{}, fmt.Errorf("creating migration job: %w", err) } - logger.Info("created migration job", "job", job.Name, "currentVersion", currentVersion, "desiredVersion", desiredVersion) + logger.Info("created migration job", "job", job.Name, "currentVersion", current.Version, "desiredVersion", desired.Version) + r.Recorder.Eventf(deployment, corev1.EventTypeNormal, "MigrationStarted", "Created migration job %s for version %s", job.Name, desired.Version) return ctrl.Result{RequeueAfter: 10 * time.Second}, nil } - if err != nil { - return ctrl.Result{}, fmt.Errorf("getting migration job: %w", err) + + if job.DeletionTimestamp != nil { + // Foreground deletion in progress: the name frees up once the pods are gone. + return ctrl.Result{RequeueAfter: 5 * time.Second}, nil + } + if !replaceable(job, deployment) { + r.Recorder.Eventf(deployment, corev1.EventTypeWarning, "MigrationJobConflict", "Job %s exists but was not created by the operator or the chart; delete or rename it to let the migration run", job.Name) + return ctrl.Result{}, fmt.Errorf("job %s/%s is not managed by the operator", job.Namespace, job.Name) } - // Only a Job this operator created for the desired version is trusted. + // Only a Job this operator created for the desired identity is trusted. // Anything else under the same name, such as a Job for a previous image or // the chart's legacy Helm hook Job, is replaced. - jobVersion := job.Annotations[AnnotationDesiredVersion] complete := isJobConditionTrue(job, batchv1.JobComplete) failedAt, failed := jobFailedAt(job) running := !complete && !failed && ptr.Deref(job.Status.Ready, 0) > 0 - outdated := jobVersion != desiredVersion + outdated := jobIdentity(job) != desired // A Job whose pod cannot start (a bad secret reference, an image pull // error, an unschedulable pod) never fails on its own, so rebuild it once // the Deployment's pod template has changed. A pod that has just finished // may still be rebuilt, which only re-runs a no-op migration. if !outdated && !complete && !failed && !running && job.Status.Active > 0 { - want := r.buildMigrationJob(deployment, container, desiredVersion) + want, err := r.buildMigrationJob(deployment, container, desired) + if err != nil { + return ctrl.Result{}, err + } outdated = job.Annotations[AnnotationPodTemplateHash] != want.Annotations[AnnotationPodTemplateHash] } if outdated { @@ -103,10 +130,10 @@ func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) ( // a concurrent index build that is aborted halfway leaves the schema in // a state the next run does not repair. Replace the Job once it ends. if running { - logger.V(1).Info("waiting for running migration job before replacing it", "job", job.Name, "jobVersion", jobVersion, "desiredVersion", desiredVersion) + logger.V(1).Info("waiting for running migration job before replacing it", "job", job.Name, "jobVersion", jobIdentity(job).Version, "desiredVersion", desired.Version) return ctrl.Result{RequeueAfter: 10 * time.Second}, nil } - logger.Info("replacing migration job", "job", job.Name, "jobVersion", jobVersion, "desiredVersion", desiredVersion) + logger.Info("replacing migration job", "job", job.Name, "jobVersion", jobIdentity(job).Version, "desiredVersion", desired.Version) if err := r.deleteJob(ctx, job); err != nil { return ctrl.Result{}, err } @@ -114,30 +141,37 @@ func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) ( } if complete { - if err := updateMigrationStatus(ctx, r.Client, deployment, desiredVersion, job.Name); err != nil { + if err := updateMigrationStatus(ctx, r.Client, deployment, desired, job.Name); err != nil { + return ctrl.Result{}, err + } + logger.Info("migration succeeded", "version", desired.Version) + r.Recorder.Eventf(deployment, corev1.EventTypeNormal, "MigrationSucceeded", "Database migrated to version %s", desired.Version) + // Only now: a TTL of 0 would otherwise remove the Job before its + // outcome is recorded, and a new Job would run the migration again. + if err := r.applyTTL(ctx, job); err != nil { return ctrl.Result{}, err } - logger.Info("migration succeeded", "version", desiredVersion) _, err := r.patchCondition(ctx, deployment, clearMigrationFailedCondition) return ctrl.Result{}, err } if failed { changed, err := r.patchCondition(ctx, deployment, func(d *appsv1.Deployment) bool { - return setMigrationFailedCondition(d, desiredVersion) + return setMigrationFailedCondition(d, desired.Version) }) if err != nil { return ctrl.Result{}, err } if changed { - logger.Info("migration job failed", "job", job.Name, "version", desiredVersion, "retryIn", retryDelay) + logger.Info("migration job failed", "job", job.Name, "version", desired.Version, "retryIn", retryDelay) + r.Recorder.Eventf(deployment, corev1.EventTypeWarning, "MigrationFailed", "Migration job %s failed for version %s; retrying in %s", job.Name, desired.Version, retryDelay) } // The failed Job itself is the retry timer, so the delay survives // operator restarts and leaves the pod logs around to inspect. if wait := retryDelay - time.Since(failedAt); wait > 0 { return ctrl.Result{RequeueAfter: wait}, nil } - logger.Info("retrying migration", "job", job.Name, "version", desiredVersion) + logger.Info("retrying migration", "job", job.Name, "version", desired.Version) if err := r.deleteJob(ctx, job); err != nil { return ctrl.Result{}, err } @@ -147,14 +181,38 @@ func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) ( return ctrl.Result{RequeueAfter: 10 * time.Second}, nil } +// deleteJob removes a migration Job and waits for its pods to be gone before +// the name is reused, so two migrations never run at once. The preconditions +// make sure the Job is still the one that was inspected. func (r *MigrationReconciler) deleteJob(ctx context.Context, job *batchv1.Job) error { - err := r.Delete(ctx, job, client.PropagationPolicy(metav1.DeletePropagationBackground)) - if client.IgnoreNotFound(err) != nil { + err := r.Delete(ctx, job, + client.PropagationPolicy(metav1.DeletePropagationForeground), + client.Preconditions{UID: &job.UID, ResourceVersion: &job.ResourceVersion}) + if apierrors.IsNotFound(err) || apierrors.IsConflict(err) { + // Gone already, or changed since it was read: the requeue re-evaluates it. + return nil + } + if err != nil { return fmt.Errorf("deleting migration job %s: %w", job.Name, err) } return nil } +// applyTTL sets ttlSecondsAfterFinished on a completed Job so Kubernetes +// cleans it up. Failed Jobs never get a TTL: the operator keeps them for the +// retry delay and deletes them itself. +func (r *MigrationReconciler) applyTTL(ctx context.Context, job *batchv1.Job) error { + if ptr.Deref(job.Spec.TTLSecondsAfterFinished, -1) == r.TTLSecondsAfterFinished { + return nil + } + patch := client.MergeFrom(job.DeepCopy()) + job.Spec.TTLSecondsAfterFinished = ptr.To(r.TTLSecondsAfterFinished) + if err := r.Patch(ctx, job, patch); err != nil { + return fmt.Errorf("setting ttlSecondsAfterFinished on job %s: %w", job.Name, client.IgnoreNotFound(err)) + } + return nil +} + // patchCondition applies update to the Deployment's status conditions and // patches the status only if something changed. The strategic merge patch // merges conditions by type, so the Deployment controller's own conditions are diff --git a/operator/internal/controller/migration_controller_test.go b/operator/internal/controller/migration_controller_test.go index 34cc39d2..15ca4d15 100644 --- a/operator/internal/controller/migration_controller_test.go +++ b/operator/internal/controller/migration_controller_test.go @@ -3,6 +3,7 @@ package controller import ( "context" "fmt" + "strings" "testing" "time" @@ -14,6 +15,7 @@ import ( "k8s.io/apimachinery/pkg/runtime" "k8s.io/apimachinery/pkg/types" clientgoscheme "k8s.io/client-go/kubernetes/scheme" + "k8s.io/client-go/tools/record" "k8s.io/utils/ptr" ctrl "sigs.k8s.io/controller-runtime" "sigs.k8s.io/controller-runtime/pkg/client" @@ -66,7 +68,10 @@ func newTestDeployment(image string) *appsv1.Deployment { // newTestJob returns the migration Job the operator builds for dep. func newTestJob(dep *appsv1.Deployment, conditions ...batchv1.JobCondition) *batchv1.Job { container := &dep.Spec.Template.Spec.Containers[0] - job := (&MigrationReconciler{}).buildMigrationJob(dep, container, extractImageTag(container.Image)) + job, err := (&MigrationReconciler{TTLSecondsAfterFinished: DefaultTTLSecondsAfterFinished}).buildMigrationJob(dep, container, desiredIdentity(dep, container)) + if err != nil { + panic(err) + } job.Status.Conditions = conditions return job } @@ -75,10 +80,15 @@ func jobCondition(t batchv1.JobConditionType, at time.Time) batchv1.JobCondition return batchv1.JobCondition{Type: t, Status: corev1.ConditionTrue, LastTransitionTime: metav1.NewTime(at)} } +// newStatus returns a status ConfigMap owned by the test Deployment. func newStatus(version string) *corev1.ConfigMap { return &corev1.ConfigMap{ - ObjectMeta: metav1.ObjectMeta{Name: statusKey.Name, Namespace: statusKey.Namespace}, - Data: map[string]string{"version": version}, + ObjectMeta: metav1.ObjectMeta{ + Name: statusKey.Name, + Namespace: statusKey.Namespace, + OwnerReferences: []metav1.OwnerReference{ownerReference(newTestDeployment(""))}, + }, + Data: map[string]string{"version": version}, } } @@ -96,12 +106,25 @@ func newReconciler(t *testing.T, funcs *interceptor.Funcs, objects ...client.Obj } return &MigrationReconciler{ Client: b.Build(), + Recorder: record.NewFakeRecorder(20), BackoffLimit: DefaultBackoffLimit, ActiveDeadlineSeconds: DefaultActiveDeadlineSeconds, TTLSecondsAfterFinished: DefaultTTLSecondsAfterFinished, } } +func events(r *MigrationReconciler) []string { + var out []string + for { + select { + case e := <-r.Recorder.(*record.FakeRecorder).Events: + out = append(out, e) + default: + return out + } + } +} + var failStatusPatch = &interceptor.Funcs{ SubResourcePatch: func(ctx context.Context, c client.Client, subResource string, obj client.Object, patch client.Patch, opts ...client.SubResourcePatchOption) error { if subResource == "status" { @@ -181,6 +204,12 @@ func TestReconcile_FirstInstall_CreatesJob(t *testing.T) { if job.Spec.ActiveDeadlineSeconds != nil { t.Errorf("expected no deadline by default, got %d", *job.Spec.ActiveDeadlineSeconds) } + if job.Spec.TTLSecondsAfterFinished != nil { + t.Errorf("TTL must only be applied after success, got %d", *job.Spec.TTLSecondsAfterFinished) + } + if len(job.Spec.Template.Spec.InitContainers) != 0 { + t.Errorf("expected no sidecars for a single-container deployment, got %d", len(job.Spec.Template.Spec.InitContainers)) + } if len(job.OwnerReferences) != 1 || !ptr.Deref(job.OwnerReferences[0].Controller, false) || job.OwnerReferences[0].BlockOwnerDeletion != nil { t.Errorf("expected a single controller owner reference without blockOwnerDeletion, got %+v", job.OwnerReferences) } @@ -270,6 +299,58 @@ func TestReconcile_JobSucceeded_CreatesStatus(t *testing.T) { if cond == nil || cond.Status != corev1.ConditionFalse || cond.Reason != "MigrationSucceeded" { t.Errorf("expected MigrationFailed=False/MigrationSucceeded, got %+v", cond) } + job, err := getJob(r) + if err != nil { + t.Fatalf("the completed job is left for Kubernetes to clean up: %v", err) + } + if got := ptr.Deref(job.Spec.TTLSecondsAfterFinished, -1); got != DefaultTTLSecondsAfterFinished { + t.Errorf("expected TTL %d applied after success, got %d", DefaultTTLSecondsAfterFinished, got) + } +} + +func TestReconcile_TTLZero_AppliedOnlyAfterSuccess(t *testing.T) { + dep := newTestDeployment("openfga/openfga:v1.14.0") + failedJob := newTestJob(dep, jobCondition(batchv1.JobFailed, time.Now())) + r := newReconciler(t, nil, dep, failedJob) + r.TTLSecondsAfterFinished = 0 + + reconcileOnce(t, r) + job, err := getJob(r) + if err != nil { + t.Fatalf("failed job must be kept for the retry delay: %v", err) + } + if job.Spec.TTLSecondsAfterFinished != nil { + t.Errorf("a failed job must not get a TTL, got %d", *job.Spec.TTLSecondsAfterFinished) + } + + job.Status = batchv1.JobStatus{Succeeded: 1, Conditions: []batchv1.JobCondition{jobCondition(batchv1.JobComplete, time.Now())}} + if err := r.Status().Update(context.Background(), job); err != nil { + t.Fatal(err) + } + reconcileOnce(t, r) + if _, err := getStatus(r); err != nil { + t.Fatalf("status must be recorded before the job can be garbage collected: %v", err) + } + job, _ = getJob(r) + if got := ptr.Deref(job.Spec.TTLSecondsAfterFinished, -1); got != 0 { + t.Errorf("expected TTL 0 after success, got %d", got) + } +} + +func TestReconcile_UpToDate_AppliesMissingTTL(t *testing.T) { + // The TTL patch failed after the status was recorded; the up-to-date path + // must finish the job off so it does not linger forever. + dep := newTestDeployment("openfga/openfga:v1.14.0") + r := newReconciler(t, nil, dep, newStatus("v1.14.0"), newTestJob(dep, jobCondition(batchv1.JobComplete, time.Now()))) + + reconcileOnce(t, r) + job, err := getJob(r) + if err != nil { + t.Fatal(err) + } + if got := ptr.Deref(job.Spec.TTLSecondsAfterFinished, -1); got != DefaultTTLSecondsAfterFinished { + t.Errorf("expected TTL %d, got %d", DefaultTTLSecondsAfterFinished, got) + } } func TestReconcile_JobSucceeded_UpdatesStatus(t *testing.T) { @@ -360,12 +441,26 @@ func TestReconcile_JobFailed_StatusPatchErrorKeepsJob(t *testing.T) { } func TestReconcile_JobForOtherVersion_Replaced(t *testing.T) { - r := newReconciler(t, nil, newTestDeployment("openfga/openfga:v1.15.0"), - newTestJob(newTestDeployment("openfga/openfga:v1.14.0"), jobCondition(batchv1.JobComplete, time.Now()))) + stale := newTestJob(newTestDeployment("openfga/openfga:v1.14.0"), jobCondition(batchv1.JobComplete, time.Now())) + var deleteOpts client.DeleteOptions + r := newReconciler(t, &interceptor.Funcs{ + Delete: func(ctx context.Context, c client.WithWatch, obj client.Object, opts ...client.DeleteOption) error { + deleteOpts.ApplyOptions(opts) + return c.Delete(ctx, obj, opts...) + }, + }, newTestDeployment("openfga/openfga:v1.15.0"), stale) if result := reconcileOnce(t, r); result.RequeueAfter == 0 { t.Error("expected a requeue after deleting the stale job") } + // Foreground deletion keeps the name taken until the old pods are gone, so + // two migrations never overlap; the preconditions pin the inspected object. + if ptr.Deref(deleteOpts.PropagationPolicy, "") != metav1.DeletePropagationForeground { + t.Errorf("expected foreground deletion, got %v", deleteOpts.PropagationPolicy) + } + if deleteOpts.Preconditions == nil || ptr.Deref(deleteOpts.Preconditions.UID, "") != stale.UID || deleteOpts.Preconditions.ResourceVersion == nil { + t.Errorf("expected UID and resourceVersion preconditions, got %+v", deleteOpts.Preconditions) + } if _, err := getJob(r); !apierrors.IsNotFound(err) { t.Errorf("expected the stale job to be deleted, got err=%v", err) } @@ -466,6 +561,144 @@ func TestReconcile_RunningJobForOtherVersion_KeptUntilItEnds(t *testing.T) { } } +func TestReconcile_TriggerChange_RunsMigrationAgain(t *testing.T) { + // Same image, but the chart's datastore configuration changed (or the user + // bumped migration.trigger): the recorded identity no longer matches. + dep := newTestDeployment("openfga/openfga:v1.14.0") + dep.Annotations[AnnotationMigrationTrigger] = "b" + status := newStatus("v1.14.0") + status.Data["trigger"] = "a" + r := newReconciler(t, nil, dep, status) + + reconcileOnce(t, r) + job, err := getJob(r) + if err != nil { + t.Fatalf("expected a migration job for the new trigger: %v", err) + } + if job.Annotations[AnnotationMigrationTrigger] != "b" { + t.Errorf("expected the job to carry trigger b, got %q", job.Annotations[AnnotationMigrationTrigger]) + } + + job.Status = batchv1.JobStatus{Succeeded: 1, Conditions: []batchv1.JobCondition{jobCondition(batchv1.JobComplete, time.Now())}} + if err := r.Status().Update(context.Background(), job); err != nil { + t.Fatal(err) + } + reconcileOnce(t, r) + cm, _ := getStatus(r) + if cm.Data["version"] != "v1.14.0" || cm.Data["trigger"] != "b" { + t.Errorf("expected version v1.14.0 and trigger b recorded, got %v", cm.Data) + } + + // A running job for the old trigger is stale and gets replaced once it ends. + if got := jobIdentity(newTestJob(dep)); got != desiredIdentity(dep, &dep.Spec.Template.Spec.Containers[0]) { + t.Errorf("job identity %+v must match the desired identity", got) + } +} + +func TestReconcile_UnownedStatus_NotTrusted(t *testing.T) { + // A ConfigMap with the status name that this Deployment does not own (a + // hand-made one, or left by a Deployment with another UID) says nothing + // about this database. Migrate, then take the ConfigMap over. + dep := newTestDeployment("openfga/openfga:v1.14.0") + foreign := newStatus("v1.14.0") + foreign.OwnerReferences = nil + r := newReconciler(t, nil, dep, foreign, newTestJob(dep, jobCondition(batchv1.JobComplete, time.Now()))) + + reconcileOnce(t, r) + cm, err := getStatus(r) + if err != nil { + t.Fatal(err) + } + if !metav1.IsControlledBy(cm, dep) { + t.Errorf("expected the status ConfigMap to be owned by the Deployment after recording, got %+v", cm.OwnerReferences) + } +} + +func TestReconcile_UnownedJob_NotTouched(t *testing.T) { + foreign := &batchv1.Job{ + ObjectMeta: metav1.ObjectMeta{Name: jobKey.Name, Namespace: jobKey.Namespace, Labels: map[string]string{"app": "someone-else"}}, + Status: batchv1.JobStatus{Conditions: []batchv1.JobCondition{jobCondition(batchv1.JobComplete, time.Now())}}, + } + r := newReconciler(t, nil, newTestDeployment("openfga/openfga:v1.14.0"), foreign) + + if _, err := r.Reconcile(context.Background(), ctrl.Request{NamespacedName: deploymentKey}); err == nil { + t.Fatal("expected an error for a job the operator does not manage") + } + if _, err := getJob(r); err != nil { + t.Errorf("a job owned by someone else must not be deleted: %v", err) + } + if _, err := getStatus(r); !apierrors.IsNotFound(err) { + t.Errorf("a foreign job must not be recorded as a migration, got err=%v", err) + } + if evs := events(r); len(evs) != 1 || !strings.Contains(evs[0], "MigrationJobConflict") { + t.Errorf("expected a MigrationJobConflict event, got %v", evs) + } +} + +func TestReconcile_CompletedJobFromRecreatedDeployment_Trusted(t *testing.T) { + // The Deployment was deleted and recreated (new UID) while the operator's + // completed Job survived. It carries the operator's managed-by label and + // the same identity, so its outcome is recorded rather than re-run. + dep := newTestDeployment("openfga/openfga:v1.14.0") + old := newTestJob(dep, jobCondition(batchv1.JobComplete, time.Now())) + old.OwnerReferences[0].UID = "previous-uid" + r := newReconciler(t, nil, dep, old) + + reconcileOnce(t, r) + cm, err := getStatus(r) + if err != nil { + t.Fatalf("expected the operator's own completed job to be trusted: %v", err) + } + if cm.Data["version"] != "v1.14.0" || !metav1.IsControlledBy(cm, dep) { + t.Errorf("expected v1.14.0 recorded and owned by the new Deployment, got %v %v", cm.Data, cm.OwnerReferences) + } +} + +func TestBuildMigrationJob_SidecarsAndMetadata(t *testing.T) { + dep := newTestDeployment("openfga/openfga:v1.14.0") + dep.Spec.Template.Spec.InitContainers = []corev1.Container{{Name: "wait-for-db", Image: "busybox"}} + dep.Spec.Template.Spec.Containers = append(dep.Spec.Template.Spec.Containers, + corev1.Container{Name: "cloud-sql-proxy", Image: "gcr.io/cloud-sql-connectors/cloud-sql-proxy:2"}) + dep.Annotations[AnnotationMigrationLabels] = `{"team":"auth","app.kubernetes.io/component":"hijack"}` + dep.Annotations[AnnotationMigrationAnnotations] = `{"sidecar.istio.io/inject":"false","openfga.dev/desired-version":"spoof"}` + + job := newTestJob(dep) + inits := job.Spec.Template.Spec.InitContainers + if len(inits) != 2 || inits[0].Name != "cloud-sql-proxy" || inits[1].Name != "wait-for-db" { + t.Fatalf("expected the proxy sidecar then the init container, got %+v", inits) + } + // A native sidecar is stopped when the migrate container exits, so the Job completes. + if ptr.Deref(inits[0].RestartPolicy, "") != corev1.ContainerRestartPolicyAlways { + t.Errorf("expected the sidecar to have restartPolicy Always, got %v", inits[0].RestartPolicy) + } + if inits[1].RestartPolicy != nil { + t.Errorf("init containers keep their restart policy, got %v", *inits[1].RestartPolicy) + } + if len(job.Spec.Template.Spec.Containers) != 1 || job.Spec.Template.Spec.Containers[0].Name != "migrate-database" { + t.Errorf("expected only the migrate container, got %+v", job.Spec.Template.Spec.Containers) + } + for _, meta := range []metav1.ObjectMeta{job.ObjectMeta, job.Spec.Template.ObjectMeta} { + if meta.Labels["team"] != "auth" || meta.Annotations["sidecar.istio.io/inject"] != "false" { + t.Errorf("expected user metadata to be forwarded, got labels=%v annotations=%v", meta.Labels, meta.Annotations) + } + if meta.Labels[LabelComponent] != "migration" { + t.Errorf("operator identity labels must win, got %v", meta.Labels) + } + } + if job.Annotations[AnnotationDesiredVersion] != "v1.14.0" { + t.Errorf("operator annotations must win, got %v", job.Annotations) + } +} + +func TestReconcile_InvalidMetadataAnnotation_ReturnsError(t *testing.T) { + dep := newTestDeployment("openfga/openfga:v1.14.0") + dep.Annotations[AnnotationMigrationLabels] = "not json" + r := newReconciler(t, nil, dep) + if _, err := r.Reconcile(context.Background(), ctrl.Request{NamespacedName: deploymentKey}); err == nil { + t.Fatal("expected an error for an unparsable migration-labels annotation") + } +} + // The legacy chart's Helm hook Job has the same name and carries the chart's // app.kubernetes.io/version label, which is the chart appVersion rather than // the image it ran. It must never be taken as proof of a migration. From 0697705dbf9d4e200ae87a53864b89af57223371 Mon Sep 17 00:00:00 2001 From: Siddhant Khare <siddhant@usegitai.com> Date: Wed, 23 Sep 2026 17:17:15 +0530 Subject: [PATCH 64/70] test(operator): cover release version guard --- .../scripts/check-operator-release_test.sh | 94 +++++++++++++++++++ .github/workflows/operator.yml | 5 + 2 files changed, 99 insertions(+) create mode 100755 .github/scripts/check-operator-release_test.sh diff --git a/.github/scripts/check-operator-release_test.sh b/.github/scripts/check-operator-release_test.sh new file mode 100755 index 00000000..7843e03b --- /dev/null +++ b/.github/scripts/check-operator-release_test.sh @@ -0,0 +1,94 @@ +#!/usr/bin/env bash +set -euo pipefail + +repo=$(git rev-parse --show-toplevel) +check="$repo/.github/scripts/check-operator-release.sh" +failures=0 + +run_case() { + local name=$1 expected=$2 image_changed=$3 operator_version=$4 operator_app_version=$5 + local parent_version=$6 parent_dependency=$7 lock_dependency=$8 + local case_dir base result=0 + + case_dir=$(mktemp -d) + mkdir -p "$case_dir/operator/internal/controller" "$case_dir/charts/openfga-operator" "$case_dir/charts/openfga" + ( + cd "$case_dir" + git init -q + git config user.name "Release Guard Test" + git config user.email "release-guard@example.com" + + printf 'package controller\n' > operator/internal/controller/controller.go + cat > charts/openfga-operator/Chart.yaml <<'EOF' +apiVersion: v2 +name: openfga-operator +version: "1.0.0" +appVersion: "1.0.0" +EOF + cat > charts/openfga/Chart.yaml <<'EOF' +apiVersion: v2 +name: openfga +version: "1.0.0" +dependencies: + - name: openfga-operator + version: "1.0.0" +EOF + cat > charts/openfga/Chart.lock <<'EOF' +dependencies: +- name: openfga-operator + version: 1.0.0 +EOF + git add . + git commit -qm base + base=$(git rev-parse HEAD) + + if [[ "$image_changed" == "true" ]]; then + printf 'var changed = true\n' >> operator/internal/controller/controller.go + fi + cat > charts/openfga-operator/Chart.yaml <<EOF +apiVersion: v2 +name: openfga-operator +version: "$operator_version" +appVersion: "$operator_app_version" +EOF + cat > charts/openfga/Chart.yaml <<EOF +apiVersion: v2 +name: openfga +version: "$parent_version" +dependencies: + - name: openfga-operator + version: "$parent_dependency" +EOF + cat > charts/openfga/Chart.lock <<EOF +dependencies: +- name: openfga-operator + version: $lock_dependency +EOF + git add . + if ! git diff --cached --quiet; then + git commit -qm candidate + fi + + "$check" "$base" >/dev/null 2>&1 || result=$? + if [[ "$expected" == "pass" && "$result" -ne 0 ]] || + [[ "$expected" == "fail" && "$result" -eq 0 ]]; then + printf 'FAIL: %s expected %s, exit code %d\n' "$name" "$expected" "$result" + exit 1 + fi + ) || failures=$((failures + 1)) +} + +run_case "unchanged image" pass false 1.0.0 1.0.0 1.0.0 1.0.0 1.0.0 +run_case "consistent release" pass true 1.1.0 1.1.0 1.1.0 1.1.0 1.1.0 +run_case "operator version unchanged" fail true 1.0.0 1.1.0 1.1.0 1.0.0 1.0.0 +run_case "operator appVersion unchanged" fail true 1.1.0 1.0.0 1.1.0 1.1.0 1.1.0 +run_case "parent version unchanged" fail true 1.1.0 1.1.0 1.0.0 1.1.0 1.1.0 +run_case "parent dependency mismatch" fail true 1.1.0 1.1.0 1.1.0 1.0.0 1.1.0 +run_case "lock dependency mismatch" fail true 1.1.0 1.1.0 1.1.0 1.1.0 1.0.0 + +if [[ "$failures" -ne 0 ]]; then + printf '%d release guard case(s) failed\n' "$failures" + exit 1 +fi + +echo "operator release guard matrix passed" diff --git a/.github/workflows/operator.yml b/.github/workflows/operator.yml index fd1a796c..f2dbc444 100644 --- a/.github/workflows/operator.yml +++ b/.github/workflows/operator.yml @@ -7,11 +7,13 @@ on: paths: - "operator/**" - "charts/openfga-operator/**" + - ".github/scripts/check-operator-release*.sh" - ".github/workflows/operator.yml" pull_request: paths: - "operator/**" - "charts/openfga-operator/**" + - ".github/scripts/check-operator-release*.sh" - ".github/workflows/operator.yml" workflow_dispatch: inputs: @@ -38,6 +40,9 @@ jobs: if: github.event_name == 'pull_request' run: .github/scripts/check-operator-release.sh "origin/${{ github.base_ref }}" + - name: Test operator release guard + run: .github/scripts/check-operator-release_test.sh + - name: Set up Go uses: actions/setup-go@40f1582b2485089dde7abd97c1529aa768e1baff # v5.6.0 with: From 3bc8ab5b3734186408950d215b49988d0f5f0180 Mon Sep 17 00:00:00 2001 From: Siddhant Khare <siddhant@usegitai.com> Date: Wed, 23 Sep 2026 17:32:21 +0530 Subject: [PATCH 65/70] fix(operator): close migration release gaps --- .github/scripts/check-operator-release.sh | 39 +++++++++++------- .../scripts/check-operator-release_test.sh | 26 +++++++----- .github/scripts/wait-for-operator-image.sh | 19 +++++++++ .../scripts/wait-for-operator-image_test.sh | 40 +++++++++++++++++++ .github/workflows/operator.yml | 7 ++++ .github/workflows/release.yml | 23 +++++++---- .../controller/migration_controller.go | 2 +- .../controller/migration_controller_test.go | 5 +-- 8 files changed, 126 insertions(+), 35 deletions(-) create mode 100755 .github/scripts/wait-for-operator-image.sh create mode 100755 .github/scripts/wait-for-operator-image_test.sh diff --git a/.github/scripts/check-operator-release.sh b/.github/scripts/check-operator-release.sh index df6e0ff6..3ee58ac2 100755 --- a/.github/scripts/check-operator-release.sh +++ b/.github/scripts/check-operator-release.sh @@ -1,19 +1,28 @@ #!/usr/bin/env bash -# Fails when the operator image changed but the chart versions that publish it -# did not. Usage: check-operator-release.sh <base-ref>, e.g. origin/main. +# Fails when operator image or chart inputs changed but the chart versions that +# publish them did not. Usage: check-operator-release.sh <base-ref>, e.g. origin/main. # # CI publishes ghcr.io/openfga/openfga-operator:<appVersion> once and never -# overwrites it, and chart-releaser skips chart versions that already exist. A -# change to the image is therefore only released when, in the same PR: -# 1. charts/openfga-operator/Chart.yaml bumps appVersion and version -# 2. charts/openfga/Chart.yaml bumps version and pins the new operator chart +# overwrites it, and chart-releaser skips chart versions that already exist. +# Release inputs are therefore only published when, in the same PR: +# 1. charts/openfga-operator/Chart.yaml bumps version +# 2. image changes also bump appVersion +# 3. charts/openfga/Chart.yaml bumps version and pins the new operator chart # Chart.lock has to match too; `helm dependency build` in CI fails if it does not. set -euo pipefail base=$1 image_inputs=(operator/cmd operator/internal operator/go.mod operator/go.sum operator/Dockerfile) -if git diff --quiet "$base...HEAD" -- "${image_inputs[@]}"; then - echo "operator image inputs unchanged" +image_changed=false +chart_changed=false +if ! git diff --quiet "$base...HEAD" -- "${image_inputs[@]}"; then + image_changed=true +fi +if ! git diff --quiet "$base...HEAD" -- charts/openfga-operator; then + chart_changed=true +fi +if [[ "$image_changed" == "false" && "$chart_changed" == "false" ]]; then + echo "operator release inputs unchanged" exit 0 fi @@ -28,18 +37,20 @@ lock_file=charts/openfga/Chart.lock operator_version=$(field version "$operator_chart") if at_base "$operator_chart" > /tmp/base-operator-chart.yaml; then - for f in appVersion version; do - if [[ "$(field $f /tmp/base-operator-chart.yaml)" == "$(field $f $operator_chart)" ]]; then - fail "$operator_chart" "operator image inputs changed but $f is still $(field $f $operator_chart); bump appVersion and version so the change is published" - fi - done + if [[ "$(field version /tmp/base-operator-chart.yaml)" == "$(field version "$operator_chart")" ]]; then + fail "$operator_chart" "operator release inputs changed but version is still $(field version "$operator_chart"); bump it so the chart change is published" + fi + if [[ "$image_changed" == "true" ]] && + [[ "$(field appVersion /tmp/base-operator-chart.yaml)" == "$(field appVersion "$operator_chart")" ]]; then + fail "$operator_chart" "operator image inputs changed but appVersion is still $(field appVersion "$operator_chart"); bump it so the image change is published" + fi else echo "operator chart is new in this PR" fi at_base "$parent_chart" > /tmp/base-parent-chart.yaml if [[ "$(field version /tmp/base-parent-chart.yaml)" == "$(field version $parent_chart)" ]]; then - fail "$parent_chart" "operator image inputs changed but the openfga chart version is still $(field version $parent_chart); bump it so a chart with the new operator is released" + fail "$parent_chart" "operator release inputs changed but the openfga chart version is still $(field version $parent_chart); bump it so a chart with the new operator is released" fi if [[ "$(dependency_version "$parent_chart")" != "$operator_version" ]]; then fail "$parent_chart" "openfga chart pins openfga-operator $(dependency_version "$parent_chart") but the operator chart is $operator_version; update the dependency and run helm dependency update charts/openfga" diff --git a/.github/scripts/check-operator-release_test.sh b/.github/scripts/check-operator-release_test.sh index 7843e03b..d3c38ab4 100755 --- a/.github/scripts/check-operator-release_test.sh +++ b/.github/scripts/check-operator-release_test.sh @@ -6,12 +6,12 @@ check="$repo/.github/scripts/check-operator-release.sh" failures=0 run_case() { - local name=$1 expected=$2 image_changed=$3 operator_version=$4 operator_app_version=$5 - local parent_version=$6 parent_dependency=$7 lock_dependency=$8 + local name=$1 expected=$2 image_changed=$3 template_changed=$4 operator_version=$5 operator_app_version=$6 + local parent_version=$7 parent_dependency=$8 lock_dependency=$9 local case_dir base result=0 case_dir=$(mktemp -d) - mkdir -p "$case_dir/operator/internal/controller" "$case_dir/charts/openfga-operator" "$case_dir/charts/openfga" + mkdir -p "$case_dir/operator/internal/controller" "$case_dir/charts/openfga-operator/templates" "$case_dir/charts/openfga" ( cd "$case_dir" git init -q @@ -25,6 +25,7 @@ name: openfga-operator version: "1.0.0" appVersion: "1.0.0" EOF + printf 'value: base\n' > charts/openfga-operator/templates/config.yaml cat > charts/openfga/Chart.yaml <<'EOF' apiVersion: v2 name: openfga @@ -45,6 +46,9 @@ EOF if [[ "$image_changed" == "true" ]]; then printf 'var changed = true\n' >> operator/internal/controller/controller.go fi + if [[ "$template_changed" == "true" ]]; then + printf 'value: candidate\n' > charts/openfga-operator/templates/config.yaml + fi cat > charts/openfga-operator/Chart.yaml <<EOF apiVersion: v2 name: openfga-operator @@ -78,13 +82,15 @@ EOF ) || failures=$((failures + 1)) } -run_case "unchanged image" pass false 1.0.0 1.0.0 1.0.0 1.0.0 1.0.0 -run_case "consistent release" pass true 1.1.0 1.1.0 1.1.0 1.1.0 1.1.0 -run_case "operator version unchanged" fail true 1.0.0 1.1.0 1.1.0 1.0.0 1.0.0 -run_case "operator appVersion unchanged" fail true 1.1.0 1.0.0 1.1.0 1.1.0 1.1.0 -run_case "parent version unchanged" fail true 1.1.0 1.1.0 1.0.0 1.1.0 1.1.0 -run_case "parent dependency mismatch" fail true 1.1.0 1.1.0 1.1.0 1.0.0 1.1.0 -run_case "lock dependency mismatch" fail true 1.1.0 1.1.0 1.1.0 1.1.0 1.0.0 +run_case "unchanged release inputs" pass false false 1.0.0 1.0.0 1.0.0 1.0.0 1.0.0 +run_case "consistent image release" pass true false 1.1.0 1.1.0 1.1.0 1.1.0 1.1.0 +run_case "operator version unchanged" fail true false 1.0.0 1.1.0 1.1.0 1.0.0 1.0.0 +run_case "operator appVersion unchanged" fail true false 1.1.0 1.0.0 1.1.0 1.1.0 1.1.0 +run_case "parent version unchanged" fail true false 1.1.0 1.1.0 1.0.0 1.1.0 1.1.0 +run_case "parent dependency mismatch" fail true false 1.1.0 1.1.0 1.1.0 1.0.0 1.1.0 +run_case "lock dependency mismatch" fail true false 1.1.0 1.1.0 1.1.0 1.1.0 1.0.0 +run_case "consistent chart-only release" pass false true 1.1.0 1.0.0 1.1.0 1.1.0 1.1.0 +run_case "chart-only version unchanged" fail false true 1.0.0 1.0.0 1.1.0 1.0.0 1.0.0 if [[ "$failures" -ne 0 ]]; then printf '%d release guard case(s) failed\n' "$failures" diff --git a/.github/scripts/wait-for-operator-image.sh b/.github/scripts/wait-for-operator-image.sh new file mode 100755 index 00000000..ff6b4076 --- /dev/null +++ b/.github/scripts/wait-for-operator-image.sh @@ -0,0 +1,19 @@ +#!/usr/bin/env bash +set -euo pipefail + +image=${1:?usage: wait-for-operator-image.sh <image> [attempts] [retry-seconds]} +attempts=${2:-120} +retry_seconds=${3:-10} + +for ((attempt = 1; attempt <= attempts; attempt++)); do + if docker buildx imagetools inspect "$image" >/dev/null 2>&1; then + echo "operator image is available: $image" + exit 0 + fi + if ((attempt < attempts)); then + sleep "$retry_seconds" + fi +done + +echo "::error::operator image was not published: $image" +exit 1 diff --git a/.github/scripts/wait-for-operator-image_test.sh b/.github/scripts/wait-for-operator-image_test.sh new file mode 100755 index 00000000..dadf57c5 --- /dev/null +++ b/.github/scripts/wait-for-operator-image_test.sh @@ -0,0 +1,40 @@ +#!/usr/bin/env bash +set -euo pipefail + +repo=$(git rev-parse --show-toplevel) +check="$repo/.github/scripts/wait-for-operator-image.sh" +workflow="$repo/.github/workflows/release.yml" +case_dir=$(mktemp -d) +trap 'rm -rf "$case_dir"' EXIT + +cat > "$case_dir/docker" <<'EOF' +#!/usr/bin/env bash +count=$(cat "$DOCKER_CALL_COUNT" 2>/dev/null || echo 0) +count=$((count + 1)) +echo "$count" > "$DOCKER_CALL_COUNT" +[[ "$count" -ge "${DOCKER_SUCCEED_ON:-999}" ]] +EOF +chmod +x "$case_dir/docker" + +export PATH="$case_dir:$PATH" +export DOCKER_CALL_COUNT="$case_dir/calls" +export DOCKER_SUCCEED_ON=3 +"$check" ghcr.io/openfga/openfga-operator:1.0.0 3 0 >/dev/null +[[ "$(cat "$DOCKER_CALL_COUNT")" == "3" ]] + +rm -f "$DOCKER_CALL_COUNT" +export DOCKER_SUCCEED_ON=999 +if "$check" ghcr.io/openfga/openfga-operator:1.0.0 2 0 >/dev/null; then + echo "expected a missing image to fail the release gate" + exit 1 +fi +[[ "$(cat "$DOCKER_CALL_COUNT")" == "2" ]] + +wait_line=$(grep -n 'name: Wait for matching operator image' "$workflow" | cut -d: -f1) +release_line=$(grep -n 'name: Run chart-releaser' "$workflow" | cut -d: -f1) +if [[ -z "$wait_line" || -z "$release_line" || "$wait_line" -ge "$release_line" ]]; then + echo "operator image gate must run before chart-releaser" + exit 1 +fi + +echo "operator image release gate passed" diff --git a/.github/workflows/operator.yml b/.github/workflows/operator.yml index f2dbc444..a8ba6cee 100644 --- a/.github/workflows/operator.yml +++ b/.github/workflows/operator.yml @@ -8,13 +8,17 @@ on: - "operator/**" - "charts/openfga-operator/**" - ".github/scripts/check-operator-release*.sh" + - ".github/scripts/wait-for-operator-image*.sh" - ".github/workflows/operator.yml" + - ".github/workflows/release.yml" pull_request: paths: - "operator/**" - "charts/openfga-operator/**" - ".github/scripts/check-operator-release*.sh" + - ".github/scripts/wait-for-operator-image*.sh" - ".github/workflows/operator.yml" + - ".github/workflows/release.yml" workflow_dispatch: inputs: push_image: @@ -43,6 +47,9 @@ jobs: - name: Test operator release guard run: .github/scripts/check-operator-release_test.sh + - name: Test operator image release gate + run: .github/scripts/wait-for-operator-image_test.sh + - name: Set up Go uses: actions/setup-go@40f1582b2485089dde7abd97c1529aa768e1baff # v5.6.0 with: diff --git a/.github/workflows/release.yml b/.github/workflows/release.yml index 914c9605..7a849f4a 100644 --- a/.github/workflows/release.yml +++ b/.github/workflows/release.yml @@ -42,6 +42,22 @@ jobs: helm repo add openfga https://openfga.github.io/helm-charts helm repo update + - name: Login to GHCR + uses: docker/login-action@dbcb813823bdd20940b903addbd779551569679f # v4.6.0 + with: + registry: ghcr.io + username: ${{ github.actor }} + password: ${{ secrets.GITHUB_TOKEN }} + + - name: Set up Docker Buildx + uses: docker/setup-buildx-action@8d2750c68a42422c14e847fe6c8ac0403b4cbd6f # v3.12.0 + + - name: Wait for matching operator image + run: | + version=$(awk '/^appVersion:/{gsub(/"/, "", $2); print $2}' charts/openfga-operator/Chart.yaml) + .github/scripts/wait-for-operator-image.sh \ + "ghcr.io/${{ github.repository_owner }}/openfga-operator:${version}" + - name: Run chart-releaser uses: helm/chart-releaser-action@cae68fefc6b5f367a0275617c9f83181ba54714f # v1.7.0 with: @@ -50,13 +66,6 @@ jobs: CR_TOKEN: "${{ secrets.GITHUB_TOKEN }}" CR_SKIP_EXISTING: true - - name: Login to GHCR - uses: docker/login-action@dbcb813823bdd20940b903addbd779551569679f # v4.6.0 - with: - registry: ghcr.io - username: ${{ github.actor }} - password: ${{ secrets.GITHUB_TOKEN }} - - name: Push chart to GHCR if: ${{ hashFiles('.cr-release-packages/*.tgz') != '' }} run: | diff --git a/operator/internal/controller/migration_controller.go b/operator/internal/controller/migration_controller.go index f304e824..47e6e71a 100644 --- a/operator/internal/controller/migration_controller.go +++ b/operator/internal/controller/migration_controller.go @@ -120,7 +120,7 @@ func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) ( // the chart's legacy Helm hook Job, is replaced. complete := isJobConditionTrue(job, batchv1.JobComplete) failedAt, failed := jobFailedAt(job) - started := !complete && !failed && (ptr.Deref(job.Status.Ready, 0) > 0 || job.Status.Succeeded > 0) + started := !complete && !failed && (job.Status.Active > 0 || ptr.Deref(job.Status.Ready, 0) > 0 || job.Status.Succeeded > 0) outdated := !jobOwnedByDeployment || jobIdentity(job) != desired if outdated { // Never interrupt a running migration: a non-transactional step such as diff --git a/operator/internal/controller/migration_controller_test.go b/operator/internal/controller/migration_controller_test.go index a7f84a27..fc78f5d9 100644 --- a/operator/internal/controller/migration_controller_test.go +++ b/operator/internal/controller/migration_controller_test.go @@ -506,11 +506,9 @@ func TestReconcile_JobForOtherVersion_Replaced(t *testing.T) { } } -func TestReconcile_PendingJobWithOutdatedTemplate_Replaced(t *testing.T) { +func TestReconcile_UnstartedJobWithOutdatedTemplate_Replaced(t *testing.T) { dep := newTestDeployment("openfga/openfga:v1.14.0") job := newTestJob(dep) - job.Status.Active = 1 - job.Status.Ready = ptr.To(int32(0)) dep.Spec.Template.Spec.Containers[0].Env[1].Value = "postgres://db.example.com/openfga" r := newReconciler(t, nil, dep, job) @@ -619,6 +617,7 @@ func TestReconcile_StartedJobWithOutdatedTemplate_Kept(t *testing.T) { name string status batchv1.JobStatus }{ + {"active pod not ready", batchv1.JobStatus{Active: 1}}, {"pod running", batchv1.JobStatus{Active: 1, Ready: ptr.To(int32(1))}}, {"pod finished before the job is marked complete", batchv1.JobStatus{Succeeded: 1, Ready: ptr.To(int32(0))}}, } { From a7e467c684b8471a66fe39765f902854e3344e48 Mon Sep 17 00:00:00 2001 From: SoulPancake <angbpy@gmail.com> Date: Wed, 23 Sep 2026 18:00:12 +0530 Subject: [PATCH 66/70] operator: fix stuck-Job rebuild, datastore trigger and migration identity Three regressions from ca25836 and 3bc8ab5, each reproduced on a cluster: - A Job whose pod exists but cannot start (CreateContainerConfigError, ImagePullBackOff, unschedulable) counted as started, so it was never rebuilt after the Deployment was fixed. Only a Ready or already succeeded pod counts as started. - The datastore URI was dropped from the migration trigger, so pointing a release at another database with the same image ran no migration and left the new pods NotReady. The URI is back in the trigger with the credentials removed, so a password change does not count and no secret goes into the hash. - The pod template hash was part of the recorded identity, so every pod change such as a log level ran a migration. The identity is the image and the trigger again; the hash only decides whether a Job that has not started is rebuilt. --- charts/openfga/README.md | 2 +- charts/openfga/templates/NOTES.txt | 4 +- charts/openfga/templates/_helpers.tpl | 13 ++- charts/openfga/tests/operator_mode_test.yaml | 21 ++-- docs/adr/002-operator-managed-migrations.md | 4 +- operator/README.md | 6 +- operator/internal/controller/helpers.go | 25 ++-- .../controller/migration_controller.go | 14 ++- .../controller/migration_controller_test.go | 108 +++++++++--------- 9 files changed, 107 insertions(+), 90 deletions(-) diff --git a/charts/openfga/README.md b/charts/openfga/README.md index 8b446994..53a61121 100644 --- a/charts/openfga/README.md +++ b/charts/openfga/README.md @@ -164,7 +164,7 @@ datastore: uriSecret: my-postgres-secret ``` -The operator only runs migrations; replicas, autoscaling and the pod template stay under the chart's control. It records the migrated image, migration trigger, and Job pod-template identity in the `<release>-migration-status` ConfigMap and sets a `MigrationFailed` condition on the Deployment if a migration fails. The Job inherits runtime pod scheduling and sidecars, plus migration-specific init containers, sidecars, volumes, mounts, resources, timeout, non-hook annotations, and labels from `migrate.*` and `datastore.migrations.resources`. It runs as a dedicated `<release>-migration` service account (`migration.serviceAccount`), which can carry cloud IAM annotations for DDL permissions. Set `migration.trigger` to a new value to rerun a migration after rotating referenced Secret data without changing the Secret name. See the [operator README](../../operator/README.md) for how it works and its limitations. +The operator only runs migrations; replicas, autoscaling and the pod template stay under the chart's control. It records the migrated version in the `<release>-migration-status` ConfigMap and sets a `MigrationFailed` condition on the Deployment if a migration fails. The migration Job is built from the OpenFGA pod spec, so `sidecars` such as a database proxy and `extraInitContainers` run alongside it, and the `migrate.*` values (labels, non-hook annotations such as `sidecar.istio.io/inject: "false"`, extra volumes, init containers, sidecars, timeout) and `datastore.migrations.resources` are applied to it. It runs as a dedicated `<release>-migration` service account (`migration.serviceAccount`), which can carry cloud IAM annotations for DDL permissions. Migrations run when the image tag or the datastore connection settings change, so pin `image.tag` to a release rather than a floating tag; set `migration.trigger` to any new value to run one on demand. See the [operator README](../../operator/README.md) for how it works and its limitations. ## Uninstalling the Chart diff --git a/charts/openfga/templates/NOTES.txt b/charts/openfga/templates/NOTES.txt index 2465fd7c..8e7b13fc 100644 --- a/charts/openfga/templates/NOTES.txt +++ b/charts/openfga/templates/NOTES.txt @@ -1,7 +1,7 @@ {{- if include "openfga.operatorMigrations" . }} NOTE: database migrations are run by the openfga-operator. Whenever the OpenFGA -image or migration inputs change it runs the {{ include "openfga.fullname" . }}-migrate Job and -records the migrated identity in the {{ include "openfga.fullname" . }}-migration-status ConfigMap. On a new database the +image or datastore settings change it runs the {{ include "openfga.fullname" . }}-migrate Job and +records the migrated version in the {{ include "openfga.fullname" . }}-migration-status ConfigMap. On a new database the OpenFGA pods stay NotReady until the first migration completes. If the pods do not become ready, check the operator and the migration Job: diff --git a/charts/openfga/templates/_helpers.tpl b/charts/openfga/templates/_helpers.tpl index fa218104..312ee3a4 100644 --- a/charts/openfga/templates/_helpers.tpl +++ b/charts/openfga/templates/_helpers.tpl @@ -86,14 +86,17 @@ Create the name of the migration service account to use (operator mode only) {{- end }} {{/* -Identity of the non-secret datastore configuration the operator migrates. -The Job template already covers rendered environment references. The explicit -migration.trigger handles changes the chart cannot safely expose, such as a -Secret value rotated under the same name. +Identity of the datastore the operator migrates. The operator runs the migration +again whenever this changes, e.g. when the chart points at a different database +with the same OpenFGA image. Credentials in the URI are left out, so a password +change does not count and no secret goes into the hash. migration.trigger is any +user-chosen string that forces another run, for changes the chart cannot see +such as a Secret rotated under the same name. */}} {{- define "openfga.migrationTrigger" -}} {{- $ds := .Values.datastore -}} -{{- dict "engine" $ds.engine "uriSecret" $ds.uriSecret "existingSecret" $ds.existingSecret "secretKeys" $ds.secretKeys "trigger" .Values.migration.trigger | toJson | sha256sum | trunc 16 -}} +{{- $uri := regexReplaceAll "^([a-z0-9+.-]+://)?[^@/]*@" (toString (default "" $ds.uri)) "${1}" -}} +{{- dict "engine" $ds.engine "uri" $uri "username" $ds.username "uriSecret" $ds.uriSecret "existingSecret" $ds.existingSecret "secretKeys" $ds.secretKeys "trigger" .Values.migration.trigger | toJson | sha256sum | trunc 16 -}} {{- end -}} {{/* diff --git a/charts/openfga/tests/operator_mode_test.yaml b/charts/openfga/tests/operator_mode_test.yaml index 025fb1ca..548b82ac 100644 --- a/charts/openfga/tests/operator_mode_test.yaml +++ b/charts/openfga/tests/operator_mode_test.yaml @@ -90,7 +90,7 @@ tests: - isNull: path: metadata.annotations - - it: should derive the migration trigger from non-secret datastore settings + - it: should derive the migration trigger from the datastore settings set: openfga-operator.enabled: true datastore.engine: postgres @@ -100,20 +100,27 @@ tests: path: metadata.annotations["openfga.dev/migration-trigger"] pattern: ^[0-9a-f]{16}$ - - it: should change the migration trigger when migration.trigger changes + - it: should change the migration trigger when the database or migration.trigger changes set: openfga-operator.enabled: true datastore.engine: postgres datastore.uri: postgres://b/openfga migration.trigger: "2" asserts: - - matchRegex: - path: metadata.annotations["openfga.dev/migration-trigger"] - pattern: ^[0-9a-f]{16}$ - notEqual: path: metadata.annotations["openfga.dev/migration-trigger"] - # value for the previous case with no explicit trigger - value: 8162fda31ac344d7 + # value for postgres://a/openfga with no explicit trigger, from the case above + value: 791636a2daefd486 + + - it: should not change the migration trigger when only the database password changes + set: + openfga-operator.enabled: true + datastore.engine: postgres + datastore.uri: postgres://openfga:other-password@a/openfga + asserts: + - equal: + path: metadata.annotations["openfga.dev/migration-trigger"] + value: 791636a2daefd486 - it: should forward migrate labels and non-hook annotations to the operator set: diff --git a/docs/adr/002-operator-managed-migrations.md b/docs/adr/002-operator-managed-migrations.md index ab8d5bbc..a60f1085 100644 --- a/docs/adr/002-operator-managed-migrations.md +++ b/docs/adr/002-operator-managed-migrations.md @@ -79,7 +79,7 @@ The operator runs a **migration controller** that reconciles the OpenFGA Deploym │ │ │ 1. Read Deployment and derive migration identity │ │ 2. Read ConfigMap/openfga-migration-status │ -│ └── "Last migrated image and pod template hash" │ +│ └── "Last migrated image and datastore trigger" │ │ 3. Identities differ → migration needed │ │ 4. Create Job/openfga-migrate │ │ ├── ServiceAccount: openfga-migrator (DDL perms) │ @@ -103,7 +103,7 @@ Readiness comes from OpenFGA itself: `IsReady()` reports `NOT_SERVING` while the #### Migration identity tracking via ConfigMap -A ConfigMap (`openfga-migration-status`) records the last successfully migrated image version, migration trigger, and Job pod-template hash. The operator compares these values to the desired Job, so changes to datastore configuration, volumes, scheduling, init containers, sidecars, and other migration inputs trigger a new migration. Because Secret contents are not present in a Deployment, users can change `migration.trigger` to force a migration after rotating a referenced Secret in place. This is: +A ConfigMap (`openfga-migration-status`) records the last successfully migrated identity: the image version and the datastore trigger the chart derives from the connection settings (without credentials). The operator compares this to the Deployment to determine if migration is needed. This is: - Simple to inspect (`kubectl get configmap openfga-migration-status -o yaml`) - Survives operator restarts - Can be manually deleted to force re-migration once the previous migration Job has been cleaned up diff --git a/operator/README.md b/operator/README.md index acfbc498..46bd0629 100644 --- a/operator/README.md +++ b/operator/README.md @@ -1,13 +1,13 @@ # OpenFGA Operator -A Kubernetes operator that manages database migrations for OpenFGA deployments. Instead of relying on Helm hooks and init containers, the operator watches OpenFGA Deployments, detects migration input changes, and orchestrates migrations as regular Jobs. +A Kubernetes operator that manages database migrations for OpenFGA deployments. Instead of relying on Helm hooks and init containers, the operator watches OpenFGA Deployments, detects image and datastore changes, and orchestrates migrations as regular Jobs. This is **Stage 1** of the operator — focused solely on migration orchestration. See [ADR-001](../docs/adr/001-adopt-openfga-operator.md) for the full roadmap. ## How It Works 1. The operator watches Deployments in its configured namespace, which defaults to the operator pod's namespace, labeled `app.kubernetes.io/part-of: openfga` and `app.kubernetes.io/component: authorization-controller` -2. When the desired migration identity changes (the image, migration trigger, or rendered Job pod template differs from the `{name}-migration-status` ConfigMap), the operator: +2. When the migration identity changes (the container image tag plus the `openfga.dev/migration-trigger` annotation, compared to the `{name}-migration-status` ConfigMap), the operator: - Creates a migration Job running `openfga migrate`, using the Deployment's image, environment, pod scheduling, init containers, and other containers as [native sidecars](https://kubernetes.io/docs/concepts/workloads/pods/sidecar-containers/) - Applies migration-specific containers, volumes, mounts, resources, timeout, labels, and annotations from the Deployment annotations rendered by the chart - Waits for the Job to complete @@ -151,7 +151,7 @@ The operator reads these annotations from the OpenFGA Deployment: ## Limitations -- **Secret contents are not observable:** The migration identity covers the image, rendered datastore settings, environment references, and pod configuration. Kubernetes does not expose referenced Secret contents through the Deployment, so change `migration.trigger` when rotating a Secret in place and a migration must rerun. +- **Secret contents are not observable:** The trigger covers the datastore settings the chart renders (engine, database host and name, Secret names and keys) but not the contents of a Secret that changes under the same name; set `migration.trigger` to a new value in that case. Credentials in the URI are never part of the trigger. - **Mutable image contents are not observable:** Reusing a tag such as `latest` does not change the Deployment's image reference. Use immutable tags or digests, or change `migration.trigger` when deliberately replacing the contents of a mutable tag. - **Helm hook metadata:** Operator-managed Jobs ignore `helm.sh/*` entries in `migrate.annotations`. Other migration annotations and labels are forwarded, but cannot override the operator's identity labels. - **Injected sidecars:** Containers injected by a webhook are not part of the Deployment's pod spec and cannot be converted to native sidecars. Disable injection for the migration pod with `migrate.annotations` if the injected container does not exit. diff --git a/operator/internal/controller/helpers.go b/operator/internal/controller/helpers.go index 31c6a0c1..410e6166 100644 --- a/operator/internal/controller/helpers.go +++ b/operator/internal/controller/helpers.go @@ -59,9 +59,8 @@ const ( // migration has to run: the OpenFGA image version plus the trigger the chart // derives from the datastore configuration. type migrationIdentity struct { - Version string - Trigger string - PodTemplateHash string + Version string + Trigger string } func desiredIdentity(deployment *appsv1.Deployment, container *corev1.Container) migrationIdentity { @@ -73,9 +72,8 @@ func desiredIdentity(deployment *appsv1.Deployment, container *corev1.Container) func jobIdentity(job *batchv1.Job) migrationIdentity { return migrationIdentity{ - Version: job.Annotations[AnnotationDesiredVersion], - Trigger: job.Annotations[AnnotationMigrationTrigger], - PodTemplateHash: job.Annotations[AnnotationPodTemplateHash], + Version: job.Annotations[AnnotationDesiredVersion], + Trigger: job.Annotations[AnnotationMigrationTrigger], } } @@ -85,11 +83,7 @@ func recordedIdentity(cm *corev1.ConfigMap, deployment *appsv1.Deployment) migra if !metav1.IsControlledBy(cm, deployment) { return migrationIdentity{} } - return migrationIdentity{ - Version: cm.Data["version"], - Trigger: cm.Data["trigger"], - PodTemplateHash: cm.Data["podTemplateHash"], - } + return migrationIdentity{Version: cm.Data["version"], Trigger: cm.Data["trigger"]} } // extractImageTag returns the tag portion of a container image reference. @@ -421,11 +415,10 @@ func updateMigrationStatus(ctx context.Context, c client.Client, deployment *app } cm.OwnerReferences = []metav1.OwnerReference{ownerReference(deployment)} cm.Data = map[string]string{ - "version": identity.Version, - "trigger": identity.Trigger, - "podTemplateHash": identity.PodTemplateHash, - "migratedAt": time.Now().UTC().Format(time.RFC3339), - "jobName": jobName, + "version": identity.Version, + "trigger": identity.Trigger, + "migratedAt": time.Now().UTC().Format(time.RFC3339), + "jobName": jobName, } return nil }) diff --git a/operator/internal/controller/migration_controller.go b/operator/internal/controller/migration_controller.go index 47e6e71a..e0c0b1b1 100644 --- a/operator/internal/controller/migration_controller.go +++ b/operator/internal/controller/migration_controller.go @@ -24,7 +24,7 @@ import ( const retryDelay = 60 * time.Second // MigrationReconciler watches OpenFGA Deployments and runs a database -// migration Job whenever its image or migration inputs change. +// migration Job whenever the OpenFGA image or the datastore trigger changes. type MigrationReconciler struct { client.Client Recorder record.EventRecorder @@ -58,7 +58,6 @@ func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) ( if err != nil { return ctrl.Result{}, err } - desired.PodTemplateHash = desiredJob.Annotations[AnnotationPodTemplateHash] status := &corev1.ConfigMap{} err = r.Get(ctx, types.NamespacedName{Name: migrationConfigMapName(req.Name), Namespace: req.Namespace}, status) @@ -120,8 +119,17 @@ func (r *MigrationReconciler) Reconcile(ctx context.Context, req ctrl.Request) ( // the chart's legacy Helm hook Job, is replaced. complete := isJobConditionTrue(job, batchv1.JobComplete) failedAt, failed := jobFailedAt(job) - started := !complete && !failed && (job.Status.Active > 0 || ptr.Deref(job.Status.Ready, 0) > 0 || job.Status.Succeeded > 0) + // A pod that is Ready is running the migration; one that already succeeded + // finished before the Job condition was written. A pod that merely exists + // (Active) may be stuck in Pending and has not started anything. + started := !complete && !failed && (ptr.Deref(job.Status.Ready, 0) > 0 || job.Status.Succeeded > 0) outdated := !jobOwnedByDeployment || jobIdentity(job) != desired + // A Job whose pod cannot start (a bad secret reference, an image pull + // error, an unschedulable pod) never fails on its own, so rebuild it once + // the Deployment's pod template has changed. + if !outdated && !complete && !failed && !started { + outdated = job.Annotations[AnnotationPodTemplateHash] != desiredJob.Annotations[AnnotationPodTemplateHash] + } if outdated { // Never interrupt a running migration: a non-transactional step such as // a concurrent index build that is aborted halfway leaves the schema in diff --git a/operator/internal/controller/migration_controller_test.go b/operator/internal/controller/migration_controller_test.go index fc78f5d9..0f10d3d0 100644 --- a/operator/internal/controller/migration_controller_test.go +++ b/operator/internal/controller/migration_controller_test.go @@ -93,9 +93,8 @@ func newStatus(dep *appsv1.Deployment) *corev1.ConfigMap { }, }, Data: map[string]string{ - "version": job.Annotations[AnnotationDesiredVersion], - "trigger": job.Annotations[AnnotationMigrationTrigger], - "podTemplateHash": job.Annotations[AnnotationPodTemplateHash], + "version": job.Annotations[AnnotationDesiredVersion], + "trigger": job.Annotations[AnnotationMigrationTrigger], }, } } @@ -298,7 +297,7 @@ func TestReconcile_JobSucceeded_CreatesStatus(t *testing.T) { if err != nil { t.Fatalf("expected migration status ConfigMap: %v", err) } - if cm.Data["version"] != "v1.14.0" || cm.Data["podTemplateHash"] != job.Annotations[AnnotationPodTemplateHash] || cm.Data["jobName"] != jobKey.Name { + if cm.Data["version"] != "v1.14.0" || cm.Data["jobName"] != jobKey.Name { t.Errorf("unexpected status data: %v", cm.Data) } if len(cm.OwnerReferences) != 1 || cm.OwnerReferences[0].UID != "test-uid-123" { @@ -395,9 +394,6 @@ func TestReconcile_JobSucceeded_UpdatesStatus(t *testing.T) { if cm.Data["version"] != "v1.14.0" { t.Errorf("expected version v1.14.0, got %q", cm.Data["version"]) } - if cm.Data["podTemplateHash"] != job.Annotations[AnnotationPodTemplateHash] { - t.Errorf("expected the completed Job's pod template hash, got %q", cm.Data["podTemplateHash"]) - } if cm.OwnerReferences[0].UID != "test-uid-123" { t.Errorf("expected owner reference to be reset to the current Deployment, got %+v", cm.OwnerReferences) } @@ -509,6 +505,9 @@ func TestReconcile_JobForOtherVersion_Replaced(t *testing.T) { func TestReconcile_UnstartedJobWithOutdatedTemplate_Replaced(t *testing.T) { dep := newTestDeployment("openfga/openfga:v1.14.0") job := newTestJob(dep) + // The pod exists but cannot start, e.g. CreateContainerConfigError. + job.Status.Active = 1 + job.Status.Ready = ptr.To(int32(0)) dep.Spec.Template.Spec.Containers[0].Env[1].Value = "postgres://db.example.com/openfga" r := newReconciler(t, nil, dep, job) @@ -527,59 +526,67 @@ func TestReconcile_UnstartedJobWithOutdatedTemplate_Replaced(t *testing.T) { } } -func TestReconcile_StatusWithOutdatedMigrationInputs_Reruns(t *testing.T) { - tests := []struct { - name string - change func(*appsv1.Deployment) - }{ - { - name: "datastore URI", - change: func(dep *appsv1.Deployment) { - dep.Spec.Template.Spec.Containers[0].Env[1].Value = "postgres://other.example.com/openfga" - }, - }, - { - name: "migration trigger", - change: func(dep *appsv1.Deployment) { - dep.Annotations[AnnotationMigrationTrigger] = "secret-rotation-2" - }, - }, +func TestReconcile_StatusWithOtherTrigger_Reruns(t *testing.T) { + dep := newTestDeployment("openfga/openfga:v1.14.0") + status := newStatus(dep) + dep.Annotations[AnnotationMigrationTrigger] = "secret-rotation-2" + r := newReconciler(t, nil, dep, status) + + reconcileOnce(t, r) + job, err := getJob(r) + if err != nil { + t.Fatalf("expected a changed trigger to create a Job: %v", err) } - for _, tt := range tests { - t.Run(tt.name, func(t *testing.T) { - dep := newTestDeployment("openfga/openfga:v1.14.0") - status := newStatus(dep) - tt.change(dep) - r := newReconciler(t, nil, dep, status) + if job.Annotations[AnnotationMigrationTrigger] != "secret-rotation-2" { + t.Errorf("expected the Job to carry the new trigger, got %q", job.Annotations[AnnotationMigrationTrigger]) + } +} - reconcileOnce(t, r) - job, err := getJob(r) - if err != nil { - t.Fatalf("expected changed migration inputs to create a Job: %v", err) - } - if job.Annotations[AnnotationPodTemplateHash] == status.Data["podTemplateHash"] { - t.Error("expected changed migration inputs to produce a new identity") - } - }) +func TestReconcile_UpToDate_IgnoresPodTemplateChanges(t *testing.T) { + // The recorded identity is the image and the trigger. A change to the pod + // template alone, such as a log level or resource limits, is not a reason + // to run the migration again. + dep := newTestDeployment("openfga/openfga:v1.14.0") + status := newStatus(dep) + dep.Spec.Template.Spec.Containers[0].Env[2].Value = "debug" + r := newReconciler(t, nil, dep, status) + + if result := reconcileOnce(t, r); result.RequeueAfter != 0 { + t.Errorf("expected no requeue for an up-to-date migration, got %v", result.RequeueAfter) + } + if _, err := getJob(r); !apierrors.IsNotFound(err) { + t.Errorf("a pod template change must not run a migration, got err=%v", err) } } -func TestReconcile_VersionOnlyStatus_Reruns(t *testing.T) { +func TestReconcile_VersionOnlyStatus(t *testing.T) { + // A status written before the trigger existed has no trigger key. It still + // matches a Deployment without a trigger, and mismatches one with a trigger. + versionOnly := func(dep *appsv1.Deployment) *corev1.ConfigMap { + return &corev1.ConfigMap{ + ObjectMeta: metav1.ObjectMeta{ + Name: statusKey.Name, + Namespace: statusKey.Namespace, + Labels: map[string]string{LabelManagedBy: LabelManagedByValue}, + OwnerReferences: []metav1.OwnerReference{ownerReference(dep)}, + }, + Data: map[string]string{"version": "v1.14.0"}, + } + } + dep := newTestDeployment("openfga/openfga:v1.14.0") - status := &corev1.ConfigMap{ - ObjectMeta: metav1.ObjectMeta{ - Name: statusKey.Name, - Namespace: statusKey.Namespace, - Labels: map[string]string{LabelManagedBy: LabelManagedByValue}, - OwnerReferences: []metav1.OwnerReference{ownerReference(dep)}, - }, - Data: map[string]string{"version": "v1.14.0"}, + r := newReconciler(t, nil, dep, versionOnly(dep)) + reconcileOnce(t, r) + if _, err := getJob(r); !apierrors.IsNotFound(err) { + t.Errorf("no trigger on either side must not run a migration, got err=%v", err) } - r := newReconciler(t, nil, dep, status) + dep = newTestDeployment("openfga/openfga:v1.14.0") + dep.Annotations[AnnotationMigrationTrigger] = "abc" + r = newReconciler(t, nil, dep, versionOnly(dep)) reconcileOnce(t, r) if _, err := getJob(r); err != nil { - t.Fatalf("expected legacy version-only status to be migrated to the new identity: %v", err) + t.Errorf("a Deployment with a trigger must migrate over a version-only status: %v", err) } } @@ -617,7 +624,6 @@ func TestReconcile_StartedJobWithOutdatedTemplate_Kept(t *testing.T) { name string status batchv1.JobStatus }{ - {"active pod not ready", batchv1.JobStatus{Active: 1}}, {"pod running", batchv1.JobStatus{Active: 1, Ready: ptr.To(int32(1))}}, {"pod finished before the job is marked complete", batchv1.JobStatus{Succeeded: 1, Ready: ptr.To(int32(0))}}, } { From 39a59117459071b71183e131db2509227de9dc50 Mon Sep 17 00:00:00 2001 From: SoulPancake <angbpy@gmail.com> Date: Wed, 23 Sep 2026 18:58:10 +0530 Subject: [PATCH 67/70] ci: drop the operator image release gate release.yml waited up to 20 minutes for ghcr.io/openfga/openfga-operator:<appVersion> before running chart-releaser. If the operator build failed on main, every chart in that push stayed unreleased until another merge, since workflow_dispatch never publishes the :<appVersion> tag. The chart already treats the openfga image this way: no registry check, the PR E2E installs it. test.yml builds and loads the operator image into kind, so a PR is proven installable before merge; the remaining window is the few minutes between chart publish and image push on the same commit, which kubelet retries through on its own. release.yml is back to its state on main. --- .github/scripts/wait-for-operator-image.sh | 19 --------- .../scripts/wait-for-operator-image_test.sh | 40 ------------------- .github/workflows/operator.yml | 7 ---- .github/workflows/release.yml | 23 ++++------- 4 files changed, 7 insertions(+), 82 deletions(-) delete mode 100755 .github/scripts/wait-for-operator-image.sh delete mode 100755 .github/scripts/wait-for-operator-image_test.sh diff --git a/.github/scripts/wait-for-operator-image.sh b/.github/scripts/wait-for-operator-image.sh deleted file mode 100755 index ff6b4076..00000000 --- a/.github/scripts/wait-for-operator-image.sh +++ /dev/null @@ -1,19 +0,0 @@ -#!/usr/bin/env bash -set -euo pipefail - -image=${1:?usage: wait-for-operator-image.sh <image> [attempts] [retry-seconds]} -attempts=${2:-120} -retry_seconds=${3:-10} - -for ((attempt = 1; attempt <= attempts; attempt++)); do - if docker buildx imagetools inspect "$image" >/dev/null 2>&1; then - echo "operator image is available: $image" - exit 0 - fi - if ((attempt < attempts)); then - sleep "$retry_seconds" - fi -done - -echo "::error::operator image was not published: $image" -exit 1 diff --git a/.github/scripts/wait-for-operator-image_test.sh b/.github/scripts/wait-for-operator-image_test.sh deleted file mode 100755 index dadf57c5..00000000 --- a/.github/scripts/wait-for-operator-image_test.sh +++ /dev/null @@ -1,40 +0,0 @@ -#!/usr/bin/env bash -set -euo pipefail - -repo=$(git rev-parse --show-toplevel) -check="$repo/.github/scripts/wait-for-operator-image.sh" -workflow="$repo/.github/workflows/release.yml" -case_dir=$(mktemp -d) -trap 'rm -rf "$case_dir"' EXIT - -cat > "$case_dir/docker" <<'EOF' -#!/usr/bin/env bash -count=$(cat "$DOCKER_CALL_COUNT" 2>/dev/null || echo 0) -count=$((count + 1)) -echo "$count" > "$DOCKER_CALL_COUNT" -[[ "$count" -ge "${DOCKER_SUCCEED_ON:-999}" ]] -EOF -chmod +x "$case_dir/docker" - -export PATH="$case_dir:$PATH" -export DOCKER_CALL_COUNT="$case_dir/calls" -export DOCKER_SUCCEED_ON=3 -"$check" ghcr.io/openfga/openfga-operator:1.0.0 3 0 >/dev/null -[[ "$(cat "$DOCKER_CALL_COUNT")" == "3" ]] - -rm -f "$DOCKER_CALL_COUNT" -export DOCKER_SUCCEED_ON=999 -if "$check" ghcr.io/openfga/openfga-operator:1.0.0 2 0 >/dev/null; then - echo "expected a missing image to fail the release gate" - exit 1 -fi -[[ "$(cat "$DOCKER_CALL_COUNT")" == "2" ]] - -wait_line=$(grep -n 'name: Wait for matching operator image' "$workflow" | cut -d: -f1) -release_line=$(grep -n 'name: Run chart-releaser' "$workflow" | cut -d: -f1) -if [[ -z "$wait_line" || -z "$release_line" || "$wait_line" -ge "$release_line" ]]; then - echo "operator image gate must run before chart-releaser" - exit 1 -fi - -echo "operator image release gate passed" diff --git a/.github/workflows/operator.yml b/.github/workflows/operator.yml index a8ba6cee..f2dbc444 100644 --- a/.github/workflows/operator.yml +++ b/.github/workflows/operator.yml @@ -8,17 +8,13 @@ on: - "operator/**" - "charts/openfga-operator/**" - ".github/scripts/check-operator-release*.sh" - - ".github/scripts/wait-for-operator-image*.sh" - ".github/workflows/operator.yml" - - ".github/workflows/release.yml" pull_request: paths: - "operator/**" - "charts/openfga-operator/**" - ".github/scripts/check-operator-release*.sh" - - ".github/scripts/wait-for-operator-image*.sh" - ".github/workflows/operator.yml" - - ".github/workflows/release.yml" workflow_dispatch: inputs: push_image: @@ -47,9 +43,6 @@ jobs: - name: Test operator release guard run: .github/scripts/check-operator-release_test.sh - - name: Test operator image release gate - run: .github/scripts/wait-for-operator-image_test.sh - - name: Set up Go uses: actions/setup-go@40f1582b2485089dde7abd97c1529aa768e1baff # v5.6.0 with: diff --git a/.github/workflows/release.yml b/.github/workflows/release.yml index 7a849f4a..914c9605 100644 --- a/.github/workflows/release.yml +++ b/.github/workflows/release.yml @@ -42,22 +42,6 @@ jobs: helm repo add openfga https://openfga.github.io/helm-charts helm repo update - - name: Login to GHCR - uses: docker/login-action@dbcb813823bdd20940b903addbd779551569679f # v4.6.0 - with: - registry: ghcr.io - username: ${{ github.actor }} - password: ${{ secrets.GITHUB_TOKEN }} - - - name: Set up Docker Buildx - uses: docker/setup-buildx-action@8d2750c68a42422c14e847fe6c8ac0403b4cbd6f # v3.12.0 - - - name: Wait for matching operator image - run: | - version=$(awk '/^appVersion:/{gsub(/"/, "", $2); print $2}' charts/openfga-operator/Chart.yaml) - .github/scripts/wait-for-operator-image.sh \ - "ghcr.io/${{ github.repository_owner }}/openfga-operator:${version}" - - name: Run chart-releaser uses: helm/chart-releaser-action@cae68fefc6b5f367a0275617c9f83181ba54714f # v1.7.0 with: @@ -66,6 +50,13 @@ jobs: CR_TOKEN: "${{ secrets.GITHUB_TOKEN }}" CR_SKIP_EXISTING: true + - name: Login to GHCR + uses: docker/login-action@dbcb813823bdd20940b903addbd779551569679f # v4.6.0 + with: + registry: ghcr.io + username: ${{ github.actor }} + password: ${{ secrets.GITHUB_TOKEN }} + - name: Push chart to GHCR if: ${{ hashFiles('.cr-release-packages/*.tgz') != '' }} run: | From e1c98790fe82f25a338b69a943f52d2c79f28933 Mon Sep 17 00:00:00 2001 From: SoulPancake <angbpy@gmail.com> Date: Wed, 23 Sep 2026 19:22:22 +0530 Subject: [PATCH 68/70] charts: add rbac.create to the operator chart, run migrations as the app service account by default The Helm hook Job ran as the OpenFGA service account. Creating a separate <release>-migration service account by default meant a release on IRSA or Workload Identity lost its IAM role on the migration Job when switching to the operator. The operator already falls back to the pod's service account when the annotation is absent, so migration.serviceAccount.create now defaults to false and the dedicated account is opt-in. rbac.create gates the operator's Role and RoleBinding, as the Helm RBAC guidelines recommend. The operator chart notes no longer warn about an unpublished image. --- charts/openfga-operator/templates/NOTES.txt | 3 --- charts/openfga-operator/templates/role.yaml | 2 ++ .../openfga-operator/templates/rolebinding.yaml | 2 ++ charts/openfga-operator/tests/rbac_test.yaml | 16 ++++++++++++++++ charts/openfga-operator/values.schema.json | 7 +++++++ charts/openfga-operator/values.yaml | 4 ++++ charts/openfga/README.md | 2 +- .../tests/operator_mode_serviceaccount_test.yaml | 9 +++++++++ charts/openfga/tests/operator_mode_test.yaml | 12 ++++++++++-- charts/openfga/values.schema.json | 4 ++-- charts/openfga/values.yaml | 6 +++--- docs/adr/001-adopt-openfga-operator.md | 2 +- docs/adr/002-operator-managed-migrations.md | 4 ++-- 13 files changed, 59 insertions(+), 14 deletions(-) create mode 100644 charts/openfga-operator/tests/rbac_test.yaml diff --git a/charts/openfga-operator/templates/NOTES.txt b/charts/openfga-operator/templates/NOTES.txt index 47dba3bc..cead36dc 100644 --- a/charts/openfga-operator/templates/NOTES.txt +++ b/charts/openfga-operator/templates/NOTES.txt @@ -1,8 +1,5 @@ The openfga-operator has been deployed. -NOTE: Ensure the operator image ({{ .Values.image.repository }}:{{ .Values.image.tag | default .Chart.AppVersion }}) is available in your registry. -If unavailable, the operator pod may remain in ImagePullBackOff until the image is pushed. - To check operator status: kubectl get deployment --namespace {{ include "openfga-operator.namespace" . }} {{ include "openfga-operator.fullname" . }} diff --git a/charts/openfga-operator/templates/role.yaml b/charts/openfga-operator/templates/role.yaml index 3ec0d99e..eb771c30 100644 --- a/charts/openfga-operator/templates/role.yaml +++ b/charts/openfga-operator/templates/role.yaml @@ -1,3 +1,4 @@ +{{- if .Values.rbac.create -}} apiVersion: rbac.authorization.k8s.io/v1 kind: Role metadata: @@ -24,3 +25,4 @@ rules: - apiGroups: [""] resources: ["events"] verbs: ["create", "patch"] +{{- end }} diff --git a/charts/openfga-operator/templates/rolebinding.yaml b/charts/openfga-operator/templates/rolebinding.yaml index 269bc3b6..d892c2c6 100644 --- a/charts/openfga-operator/templates/rolebinding.yaml +++ b/charts/openfga-operator/templates/rolebinding.yaml @@ -1,3 +1,4 @@ +{{- if .Values.rbac.create -}} apiVersion: rbac.authorization.k8s.io/v1 kind: RoleBinding metadata: @@ -13,3 +14,4 @@ subjects: - kind: ServiceAccount name: {{ include "openfga-operator.serviceAccountName" . }} namespace: {{ include "openfga-operator.namespace" . }} +{{- end }} diff --git a/charts/openfga-operator/tests/rbac_test.yaml b/charts/openfga-operator/tests/rbac_test.yaml new file mode 100644 index 00000000..c61de222 --- /dev/null +++ b/charts/openfga-operator/tests/rbac_test.yaml @@ -0,0 +1,16 @@ +suite: operator RBAC +templates: + - templates/role.yaml + - templates/rolebinding.yaml +tests: + - it: should create the Role and RoleBinding by default + asserts: + - hasDocuments: + count: 1 + + - it: should create nothing when rbac.create is false + set: + rbac.create: false + asserts: + - hasDocuments: + count: 0 diff --git a/charts/openfga-operator/values.schema.json b/charts/openfga-operator/values.schema.json index 0315d88b..d2889bbe 100644 --- a/charts/openfga-operator/values.schema.json +++ b/charts/openfga-operator/values.schema.json @@ -51,6 +51,13 @@ }, "additionalProperties": false }, + "rbac": { + "type": "object", + "properties": { + "create": { "type": "boolean" } + }, + "additionalProperties": false + }, "podAnnotations": { "type": "object" }, "podSecurityContext": { "type": "object" }, "securityContext": { "type": "object" }, diff --git a/charts/openfga-operator/values.yaml b/charts/openfga-operator/values.yaml index b6ba9f5c..e6aa56d6 100644 --- a/charts/openfga-operator/values.yaml +++ b/charts/openfga-operator/values.yaml @@ -29,6 +29,10 @@ serviceAccount: # If not set and create is true, a name is generated using the fullname template. name: "" +rbac: + # -- Create the Role and RoleBinding the operator needs in the watched namespace. + create: true + podAnnotations: {} podSecurityContext: diff --git a/charts/openfga/README.md b/charts/openfga/README.md index 53a61121..1b04c622 100644 --- a/charts/openfga/README.md +++ b/charts/openfga/README.md @@ -164,7 +164,7 @@ datastore: uriSecret: my-postgres-secret ``` -The operator only runs migrations; replicas, autoscaling and the pod template stay under the chart's control. It records the migrated version in the `<release>-migration-status` ConfigMap and sets a `MigrationFailed` condition on the Deployment if a migration fails. The migration Job is built from the OpenFGA pod spec, so `sidecars` such as a database proxy and `extraInitContainers` run alongside it, and the `migrate.*` values (labels, non-hook annotations such as `sidecar.istio.io/inject: "false"`, extra volumes, init containers, sidecars, timeout) and `datastore.migrations.resources` are applied to it. It runs as a dedicated `<release>-migration` service account (`migration.serviceAccount`), which can carry cloud IAM annotations for DDL permissions. Migrations run when the image tag or the datastore connection settings change, so pin `image.tag` to a release rather than a floating tag; set `migration.trigger` to any new value to run one on demand. See the [operator README](../../operator/README.md) for how it works and its limitations. +The operator only runs migrations; replicas, autoscaling and the pod template stay under the chart's control. It records the migrated version in the `<release>-migration-status` ConfigMap and sets a `MigrationFailed` condition on the Deployment if a migration fails. The migration Job is built from the OpenFGA pod spec, so `sidecars` such as a database proxy and `extraInitContainers` run alongside it, and the `migrate.*` values (labels, non-hook annotations such as `sidecar.istio.io/inject: "false"`, extra volumes, init containers, sidecars, timeout) and `datastore.migrations.resources` are applied to it. The Job runs as the OpenFGA service account; set `migration.serviceAccount.create` to give it a dedicated `<release>-migration` one, for example with cloud IAM annotations for DDL permissions. Migrations run when the image tag or the datastore connection settings change, so pin `image.tag` to a release rather than a floating tag; set `migration.trigger` to any new value to run one on demand. See the [operator README](../../operator/README.md) for how it works and its limitations. ## Uninstalling the Chart diff --git a/charts/openfga/tests/operator_mode_serviceaccount_test.yaml b/charts/openfga/tests/operator_mode_serviceaccount_test.yaml index 4d50202e..32959d61 100644 --- a/charts/openfga/tests/operator_mode_serviceaccount_test.yaml +++ b/charts/openfga/tests/operator_mode_serviceaccount_test.yaml @@ -27,6 +27,15 @@ tests: - hasDocuments: count: 1 + - it: should not render migration service account by default + set: + openfga-operator.enabled: true + datastore.engine: postgres + serviceAccount.create: true + asserts: + - hasDocuments: + count: 1 + - it: should not render migration service account when migration SA creation is disabled set: openfga-operator.enabled: true diff --git a/charts/openfga/tests/operator_mode_test.yaml b/charts/openfga/tests/operator_mode_test.yaml index 548b82ac..88a1dd94 100644 --- a/charts/openfga/tests/operator_mode_test.yaml +++ b/charts/openfga/tests/operator_mode_test.yaml @@ -14,6 +14,15 @@ tests: - equal: path: metadata.annotations["openfga.dev/container-name"] value: openfga + - isNull: + path: metadata.annotations["openfga.dev/migration-service-account"] + + - it: should set the migration service account annotation when the chart creates one + set: + openfga-operator.enabled: true + datastore.engine: postgres + migration.serviceAccount.create: true + asserts: - equal: path: metadata.annotations["openfga.dev/migration-service-account"] value: RELEASE-NAME-openfga-migration @@ -159,11 +168,10 @@ tests: path: metadata.annotations["openfga.dev/migration-service-account"] value: my-custom-sa - - it: should not set migration-service-account annotation when SA creation is disabled and no name set + - it: should not set migration-service-account annotation by default set: openfga-operator.enabled: true datastore.engine: postgres - migration.serviceAccount.create: false asserts: - isNull: path: metadata.annotations["openfga.dev/migration-service-account"] diff --git a/charts/openfga/values.schema.json b/charts/openfga/values.schema.json index aa96a507..e0348b90 100644 --- a/charts/openfga/values.schema.json +++ b/charts/openfga/values.schema.json @@ -1326,8 +1326,8 @@ "properties": { "create": { "type": "boolean", - "description": "Create a dedicated service account for migration Jobs", - "default": true + "description": "Create a dedicated service account for migration Jobs; otherwise they run as the OpenFGA service account", + "default": false }, "annotations": { "type": "object", diff --git a/charts/openfga/values.yaml b/charts/openfga/values.yaml index ce302c43..8c40374c 100644 --- a/charts/openfga/values.yaml +++ b/charts/openfga/values.yaml @@ -416,9 +416,9 @@ migration: # changed, e.g. after rotating a Secret to point at a different database. trigger: "" serviceAccount: - # -- Create a dedicated service account for migration Jobs. - # The migration Job inherits env vars (including secretKeyRef) from the OpenFGA container. - create: true + # -- Create a dedicated service account for migration Jobs. By default the + # Job runs as the OpenFGA service account, like the Helm hook Job does. + create: false # -- Annotations to add to the migration service account. # Use this to attach cloud IAM roles (e.g., eks.amazonaws.com/role-arn) for DDL permissions. annotations: {} diff --git a/docs/adr/001-adopt-openfga-operator.md b/docs/adr/001-adopt-openfga-operator.md index 35c58ddf..7918d787 100644 --- a/docs/adr/001-adopt-openfga-operator.md +++ b/docs/adr/001-adopt-openfga-operator.md @@ -81,7 +81,7 @@ Stage 1 has shipped on the `feat/operator-migration` branch. Stages 2-4 are plan - Operator packaged as a Helm subchart (`charts/openfga-operator/`) and wired into the main chart via a `condition: openfga-operator.enabled` dependency - `openfga-operator.enabled` values toggle (default `false`) that gates all operator-managed behavior - Migration reconciler (`migration_controller.go`) that runs migration Jobs when the operator is enabled -- Separate migration ServiceAccount with IAM-annotation support (`openfga.migrationServiceAccountName` helper), created when the operator is enabled +- Optional separate migration ServiceAccount with IAM-annotation support (`openfga.migrationServiceAccountName` helper, `migration.serviceAccount.create`) ### Deferred to later stages diff --git a/docs/adr/002-operator-managed-migrations.md b/docs/adr/002-operator-managed-migrations.md index a60f1085..9fd16683 100644 --- a/docs/adr/002-operator-managed-migrations.md +++ b/docs/adr/002-operator-managed-migrations.md @@ -110,7 +110,7 @@ A ConfigMap (`openfga-migration-status`) records the last successfully migrated #### Separate ServiceAccount for migrations -The chart creates a dedicated `{fullname}-migration` ServiceAccount that the operator uses for migration Jobs. Users can annotate it with cloud IAM roles that grant DDL permissions, while the runtime ServiceAccount retains only CRUD permissions. +Migration Jobs run as the OpenFGA ServiceAccount by default, as the Helm hook Job does. With `migration.serviceAccount.create` the chart creates a dedicated `{fullname}-migration` ServiceAccount instead, which can carry cloud IAM roles that grant DDL permissions while the runtime ServiceAccount keeps only CRUD permissions. #### Migration Job is a regular resource @@ -207,7 +207,7 @@ Nothing is deleted outright — every change is gated on `openfga-operator.enabl |--------------|---------| | `values.yaml`: `openfga-operator.enabled` | Toggle the operator subchart | | `values.yaml`: `openfga-operator.migrationJob.*` | Migration Job backoff, deadline, and TTL configuration | -| `values.yaml`: `migration.serviceAccount.*` | Separate ServiceAccount for migration Jobs | +| `values.yaml`: `migration.serviceAccount.*` | Optional separate ServiceAccount for migration Jobs | | `values.yaml`: `migration.trigger` | Explicit rerun trigger for referenced Secret data changes | | `values.yaml`: migration pod values | `migrate.extraInitContainers`, `migrate.sidecars`, volumes, mounts, resources, timeout, non-hook annotations, and labels are forwarded to operator Jobs | | `templates/serviceaccount.yaml`: second SA | Migration ServiceAccount | From 6f56fc580a679fcde6341e817e94c1eaa3ef5901 Mon Sep 17 00:00:00 2001 From: SoulPancake <angbpy@gmail.com> Date: Wed, 23 Sep 2026 19:53:19 +0530 Subject: [PATCH 69/70] operator: sign the image, disable metrics by default, document security scope The image build now attaches an SBOM and build provenance and signs the digest with cosign keyless, then verifies the signature in the same job, the way openfga/openfga releases do. controller-runtime served /metrics on 8080 to any pod in the cluster with no declared port and no authentication. The chart now passes --metrics-bind-address=0 unless metrics.enabled is set, which also declares a named container port. The operator README gets a Security section: namespace scope, the Role's verbs, ports, how to verify the image signature and where to report issues. The operator chart declares the Artifact Hub operator, capability and signing key annotations that the openfga chart already carries. --- .github/workflows/operator.yml | 24 +++++++++++++++ charts/openfga-operator/Chart.yaml | 5 ++++ .../templates/deployment.yaml | 6 ++++ .../openfga-operator/tests/metrics_test.yaml | 29 +++++++++++++++++++ charts/openfga-operator/values.schema.json | 7 +++++ charts/openfga-operator/values.yaml | 5 ++++ operator/README.md | 29 ++++++++++++++++++- 7 files changed, 104 insertions(+), 1 deletion(-) create mode 100644 charts/openfga-operator/tests/metrics_test.yaml diff --git a/.github/workflows/operator.yml b/.github/workflows/operator.yml index f2dbc444..9a2ae680 100644 --- a/.github/workflows/operator.yml +++ b/.github/workflows/operator.yml @@ -67,6 +67,7 @@ jobs: permissions: contents: read packages: write + id-token: write steps: - name: Checkout uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1 @@ -143,12 +144,16 @@ jobs: echo "Resolved tags: ${tags}" - name: Build and (conditionally) push + id: build uses: docker/build-push-action@10e90e3645eae34f1e60eeb005ba3a3d33f178e8 # v6.19.2 with: context: operator push: ${{ steps.policy.outputs.push }} platforms: linux/amd64,linux/arm64 tags: ${{ steps.tags.outputs.tags }} + # SBOM and build provenance are attached to the pushed image index. + sbom: ${{ steps.policy.outputs.push == 'true' }} + provenance: ${{ steps.policy.outputs.push == 'true' && 'mode=max' || 'false' }} cache-from: type=gha cache-to: type=gha,mode=max labels: | @@ -157,3 +162,22 @@ jobs: org.opencontainers.image.revision=${{ github.sha }} org.opencontainers.image.title=openfga-operator org.opencontainers.image.description=OpenFGA Kubernetes operator for migration orchestration + + - name: Install cosign + if: steps.policy.outputs.push == 'true' + uses: sigstore/cosign-installer@6f9f17788090df1f26f669e9d70d6ae9567deba6 # v4.1.2 + with: + cosign-release: "v2.6.1" + + # Keyless signature on the digest covers every tag pushed above. + - name: Sign and verify image + if: steps.policy.outputs.push == 'true' + env: + IMAGE: ${{ env.IMAGE_NAME }}@${{ steps.build.outputs.digest }} + IDENTITY: https://github.com/${{ github.repository }}/.github/workflows/operator.yml@${{ github.ref }} + run: | + cosign sign --yes "$IMAGE" + cosign verify "$IMAGE" \ + --certificate-oidc-issuer https://token.actions.githubusercontent.com \ + --certificate-identity "$IDENTITY" >/dev/null + echo "signed and verified $IMAGE" diff --git a/charts/openfga-operator/Chart.yaml b/charts/openfga-operator/Chart.yaml index 95da06ba..0a321ac8 100644 --- a/charts/openfga-operator/Chart.yaml +++ b/charts/openfga-operator/Chart.yaml @@ -17,3 +17,8 @@ sources: annotations: artifacthub.io/license: Apache-2.0 + artifacthub.io/operator: "true" + artifacthub.io/operatorCapabilities: Basic Install + artifacthub.io/signKey: | + fingerprint: 8E9B315F6C22E339959DA77B35CCF4BDC9F58F2A + url: https://openfga.github.io/helm-charts/pgp-public-key.asc diff --git a/charts/openfga-operator/templates/deployment.yaml b/charts/openfga-operator/templates/deployment.yaml index 5b83a6bd..e2f22189 100644 --- a/charts/openfga-operator/templates/deployment.yaml +++ b/charts/openfga-operator/templates/deployment.yaml @@ -43,6 +43,7 @@ spec: {{- if .Values.watchNamespace }} - --watch-namespace={{ .Values.watchNamespace }} {{- end }} + - --metrics-bind-address={{ if .Values.metrics.enabled }}:8080{{ else }}0{{ end }} - --backoff-limit={{ .Values.migrationJob.backoffLimit }} - --active-deadline-seconds={{ .Values.migrationJob.activeDeadlineSeconds }} - --ttl-seconds-after-finished={{ .Values.migrationJob.ttlSecondsAfterFinished }} @@ -55,6 +56,11 @@ spec: - name: healthz containerPort: 8081 protocol: TCP + {{- if .Values.metrics.enabled }} + - name: metrics + containerPort: 8080 + protocol: TCP + {{- end }} livenessProbe: httpGet: path: /healthz diff --git a/charts/openfga-operator/tests/metrics_test.yaml b/charts/openfga-operator/tests/metrics_test.yaml new file mode 100644 index 00000000..3d673809 --- /dev/null +++ b/charts/openfga-operator/tests/metrics_test.yaml @@ -0,0 +1,29 @@ +suite: operator metrics +templates: + - templates/deployment.yaml +tests: + - it: should disable the metrics endpoint by default + asserts: + - contains: + path: spec.template.spec.containers[0].args + content: --metrics-bind-address=0 + - notContains: + path: spec.template.spec.containers[0].ports + content: + name: metrics + containerPort: 8080 + protocol: TCP + + - it: should serve metrics on a named port when enabled + set: + metrics.enabled: true + asserts: + - contains: + path: spec.template.spec.containers[0].args + content: --metrics-bind-address=:8080 + - contains: + path: spec.template.spec.containers[0].ports + content: + name: metrics + containerPort: 8080 + protocol: TCP diff --git a/charts/openfga-operator/values.schema.json b/charts/openfga-operator/values.schema.json index d2889bbe..70f877d4 100644 --- a/charts/openfga-operator/values.schema.json +++ b/charts/openfga-operator/values.schema.json @@ -69,6 +69,13 @@ }, "additionalProperties": false }, + "metrics": { + "type": "object", + "properties": { + "enabled": { "type": "boolean" } + }, + "additionalProperties": false + }, "migrationJob": { "type": "object", "properties": { diff --git a/charts/openfga-operator/values.yaml b/charts/openfga-operator/values.yaml index e6aa56d6..614c9298 100644 --- a/charts/openfga-operator/values.yaml +++ b/charts/openfga-operator/values.yaml @@ -63,6 +63,11 @@ leaderElection: # -- Enable leader election for controller manager. enabled: true +metrics: + # -- Serve controller-runtime metrics on a `metrics` port (8080). Plain HTTP + # without authentication, so leave off unless something scrapes it. + enabled: false + migrationJob: # -- Number of pod failures before a migration Job is considered failed. backoffLimit: 3 diff --git a/operator/README.md b/operator/README.md index 46bd0629..e9e6ff49 100644 --- a/operator/README.md +++ b/operator/README.md @@ -122,7 +122,7 @@ The operator accepts the following flags: |------|---------|-------------| | `--leader-elect` | `false` | Enable leader election so only one replica actively reconciles at a time. Required when running multiple operator replicas for high availability; standby pods wait for the leader's Lease to expire before taking over. Not needed for single-replica deployments. | | `--watch-namespace` | `""` | Namespace to watch for OpenFGA Deployments. Defaults to the operator pod's own namespace (via `POD_NAMESPACE` env var). The chart binds namespaced RBAC in the configured watch namespace, so the operator may run in a different namespace when needed. | -| `--metrics-bind-address` | `:8080` | Address the Prometheus metrics endpoint binds to. Change only if the default port conflicts with other containers in the pod. | +| `--metrics-bind-address` | `:8080` | Address the Prometheus metrics endpoint binds to; `0` disables it. The chart passes `0` unless `metrics.enabled` is set, which also declares a `metrics` container port. The endpoint is plain HTTP without authentication. | | `--health-probe-bind-address` | `:8081` | Address the Kubernetes liveness and readiness probe endpoints bind to. Change only if the default port conflicts. | | `--backoff-limit` | `3` | Number of times a migration Job's pod can fail before the Job is considered failed. The operator then sets a `MigrationFailed` condition on the Deployment and replaces the Job 60 seconds after it failed. | | `--active-deadline-seconds` | `0` | Maximum wall-clock seconds a migration Job can run before Kubernetes terminates it. `0` means no deadline. A deadline cuts off long migrations, such as index builds or MySQL table rebuilds on large tables, which then start over on the next attempt. | @@ -149,6 +149,33 @@ The operator reads these annotations from the OpenFGA Deployment: | `openfga.dev/migration-annotations` | JSON map of non-Helm annotations for the migration Job and pod. Generated from `migrate.annotations`; `helm.sh/*` hook annotations are excluded. | | `openfga.dev/migration-labels` | JSON map of additional labels for the migration Job and pod. Generated from `migrate.labels`; operator identity labels take precedence. | +## Security + +The operator is namespace-scoped. It watches one namespace, runs with a Role rather than a ClusterRole, and never reads Secrets itself: the migration Job gets the OpenFGA container's environment, including `secretKeyRef` entries, and runs as the service account named in `openfga.dev/migration-service-account`, or the OpenFGA pod's service account when that annotation is absent. + +The Role the chart creates in the watch namespace: + +| Resource | Verbs | Used for | +|----------|-------|----------| +| `apps/deployments` | get, list, watch | Find opted-in OpenFGA Deployments | +| `apps/deployments/status` | patch | Set the `MigrationFailed` condition | +| `batch/jobs` | get, list, watch, create, delete, patch | Run, replace and expire migration Jobs | +| `configmaps` | get, list, watch, create, update | Record the migrated version | +| `coordination.k8s.io/leases` | get, list, watch, create, update | Leader election | +| `events` | create, patch | `MigrationStarted`, `MigrationSucceeded`, `MigrationFailed`, `MigrationJobConflict` | + +Ports: `8081` serves `/healthz` and `/readyz` for the kubelet. `8080` serves Prometheus metrics only when `metrics.enabled` is set; it has no authentication, so restrict it with a NetworkPolicy if the namespace is shared. + +Images pushed by `.github/workflows/operator.yml` carry an SBOM and build provenance and are signed with cosign keyless. To verify: + +```bash +cosign verify ghcr.io/openfga/openfga-operator:<tag> \ + --certificate-oidc-issuer https://token.actions.githubusercontent.com \ + --certificate-identity-regexp '^https://github.com/openfga/helm-charts/\.github/workflows/operator\.yml@refs/heads/' +``` + +Report vulnerabilities through the [security policy](https://github.com/openfga/helm-charts/security/policy), not in a public issue. + ## Limitations - **Secret contents are not observable:** The trigger covers the datastore settings the chart renders (engine, database host and name, Secret names and keys) but not the contents of a Secret that changes under the same name; set `migration.trigger` to a new value in that case. Credentials in the URI are never part of the trigger. From f6e353fd301dc293926d804fb401b42048eec125 Mon Sep 17 00:00:00 2001 From: SoulPancake <angbpy@gmail.com> Date: Fri, 25 Sep 2026 12:28:29 +0530 Subject: [PATCH 70/70] docs: mark both ADRs proposed, scope ADR-001 to stage 1 ADR-001 was self-marked Accepted before any maintainer review; the ADR process in docs/adr/README.md leaves that to the approving maintainers. It also decided stages 2-4 and the removal of the legacy path, which this PR does not deliver. Align the ADRs with the chart defaults: the migration Job runs as the OpenFGA service account unless migration.serviceAccount.create is set, and the operator requests 10m CPU. --- docs/adr/001-adopt-openfga-operator.md | 30 ++++++++++----------- docs/adr/002-operator-managed-migrations.md | 12 ++++----- docs/adr/README.md | 2 +- 3 files changed, 22 insertions(+), 22 deletions(-) diff --git a/docs/adr/001-adopt-openfga-operator.md b/docs/adr/001-adopt-openfga-operator.md index 7918d787..c808a47c 100644 --- a/docs/adr/001-adopt-openfga-operator.md +++ b/docs/adr/001-adopt-openfga-operator.md @@ -1,9 +1,9 @@ # ADR-001: Adopt a Kubernetes Operator for OpenFGA Lifecycle Management -- **Status:** Accepted — Stage 1 implemented +- **Status:** Proposed - **Date:** 2026-04-06 - **Deciders:** OpenFGA Helm Charts maintainers -- **Related Issues:** #211, #107, #120, #100, #95, #126, #132, #143, #144 +- **Related Issues:** #211, #107, #120, #100, #95, #126, #132, #144 ## Context @@ -51,16 +51,16 @@ The OpenFGA Helm chart currently handles all lifecycle concerns — deployment, ## Decision -We will build an **OpenFGA Kubernetes Operator** that handles: +We will build an **OpenFGA Kubernetes Operator**. This ADR decides Stage 1 only: 1. **Database migration orchestration** (Stage 1) — replacing Helm hooks, the `k8s-wait-for` init container, and shared ServiceAccount with operator-managed migration Jobs. -2. **Declarative store lifecycle management** (Stages 2-4) — exposing `FGAStore`, `FGAModel`, and `FGATuples` CRDs for GitOps-native authorization configuration. +2. **Declarative store lifecycle management** (Stages 2-4) — `FGAStore`, `FGAModel`, and `FGATuples` CRDs for GitOps-native authorization configuration. Under consideration; each stage needs its own ADR before implementation. The operator will be: - Written in Go using `controller-runtime` / kubebuilder - Distributed as a Helm subchart dependency of the main OpenFGA chart -- Optional — users who don't need it can set `openfga-operator.enabled: false` and fall back to the existing behavior +- Optional — `openfga-operator.enabled` defaults to `false`, which keeps the existing behavior Development will follow a staged approach to deliver value incrementally: @@ -73,7 +73,7 @@ Development will follow a staged approach to deliver value incrementally: ## Implementation Status -Stage 1 has shipped on the `feat/operator-migration` branch. Stages 2-4 are planned but not yet implemented. +Stage 1 is implemented alongside this ADR (openfga chart 0.4.0, openfga-operator chart 0.1.0). Stages 2-4 are not implemented. ### Delivered in Stage 1 @@ -88,30 +88,30 @@ Stage 1 has shipped on the `feat/operator-migration` branch. Stages 2-4 are plan - `FGAStore`, `FGAModel`, and `FGATuples` CRDs and their controllers - Declarative store/model/tuple lifecycle management -### Backward-compatibility path (deprecated) +### Backward-compatibility path -When `openfga-operator.enabled: false`, the chart still renders the legacy migration path: the Helm-hook migration Job, the `groundnuty/k8s-wait-for` init container, and the job-status RBAC. **This path is deprecated and will be removed in a future release** once the operator is the default and users have had time to migrate. It remains only to preserve backward compatibility during the transition. +When `openfga-operator.enabled: false` (the default), the chart still renders the legacy migration path: the Helm-hook migration Job, the `groundnuty/k8s-wait-for` init container, and the job-status RBAC. Making the operator the default and retiring this path is left to a later ADR. ## Consequences ### Positive -- **Resolves all 6 migration issues** (#211, #107, #120, #100, #95, #126) and related dependency issues (#132, #144) on the operator-enabled path -- **Removes `k8s-wait-for` from the operator-enabled path** — the unmaintained, CVE-carrying image is no longer used when `openfga-operator.enabled: true`, and will be removed from the chart entirely once the legacy path is retired +- **Resolves the migration issues** (#211, #107, #120, #100, #126) and related dependency issues (#132, #144) on the operator-enabled path; #95 is addressed by the opt-in migration ServiceAccount +- **Removes `k8s-wait-for` from the operator-enabled path** — the unmaintained, CVE-carrying image is no longer used when `openfga-operator.enabled: true`, and would leave the chart entirely if the legacy path is retired - **Enables GitOps-native authorization management** (planned, Stages 2-4) — stores, models, and tuples will become declarative Kubernetes resources that ArgoCD/FluxCD can sync -- **Enforces least-privilege** — separate ServiceAccounts for migration (DDL) and runtime (CRUD) on the operator-enabled path -- **Path to simplifying the Helm chart** — the migration Job template, init container logic, job-status RBAC, and hook annotations are conditionalized behind `openfga-operator.enabled: false` and scheduled for removal when the legacy path is retired +- **Enables least-privilege** — `migration.serviceAccount.create` gives migration Jobs (DDL) a ServiceAccount separate from the runtime (CRUD) on the operator-enabled path +- **Path to simplifying the Helm chart** — the migration Job template, init container logic, job-status RBAC, and hook annotations are conditionalized behind `openfga-operator.enabled: false` and could be removed if the legacy path is retired - **Follows Kubernetes ecosystem conventions** — operators are the standard pattern for managing stateful application lifecycle ### Negative - **New component to maintain** — the operator is a full Go project with its own release cycle, CI, testing, and CVE surface -- **Increased deployment footprint** — an additional pod running in the cluster (though resource requirements are minimal: ~50m CPU, ~64Mi memory) +- **Increased deployment footprint** — an additional pod running in the cluster (default requests 10m CPU, 64Mi memory) - **Learning curve** — contributors need to understand controller-runtime patterns to modify the operator - **CRD management complexity** (applies once Stages 2-4 land) — Helm does not upgrade or delete CRDs; users may need to apply CRD manifests separately on operator upgrades -- **Two code paths during the transition** — the chart must maintain both the operator-enabled path and the deprecated legacy path until the latter is removed +- **Two code paths** — the chart must maintain both the operator-enabled path and the legacy path ### Neutral -- **Backward compatibility preserved during the transition** — `openfga-operator.enabled: false` keeps the existing Helm-hook behavior working for users who have not yet migrated, but this path is deprecated and slated for removal +- **Backward compatibility preserved** — `openfga-operator.enabled: false` (the default) keeps the existing Helm-hook behavior unchanged - **No change for memory-datastore users** — users running with `datastore.engine: memory` are unaffected (no migrations, no operator needed) diff --git a/docs/adr/002-operator-managed-migrations.md b/docs/adr/002-operator-managed-migrations.md index 9fd16683..baa5a015 100644 --- a/docs/adr/002-operator-managed-migrations.md +++ b/docs/adr/002-operator-managed-migrations.md @@ -82,7 +82,7 @@ The operator runs a **migration controller** that reconciles the OpenFGA Deploym │ └── "Last migrated image and datastore trigger" │ │ 3. Identities differ → migration needed │ │ 4. Create Job/openfga-migrate │ -│ ├── ServiceAccount: openfga-migrator (DDL perms) │ +│ ├── ServiceAccount: openfga (or a dedicated one) │ │ ├── Image: openfga/openfga:v1.14.0 │ │ ├── Args: ["migrate"] │ │ └── ttlSecondsAfterFinished: 300 │ @@ -152,7 +152,7 @@ Problems: ArgoCD skips step 4. FluxCD deletes Job in step 4. `--wait` deadlocks ```text helm install - ├── Create ServiceAccount (runtime), ServiceAccount (migrator) + ├── Create ServiceAccount (plus a migration ServiceAccount if enabled) ├── Create Secret, Service ├── Create Deployment (no init containers) ├── Create Operator Deployment @@ -164,7 +164,7 @@ Operator starts: ├── Detects Deployment image version ├── No migration status ConfigMap → migration needed ├── Creates Job/openfga-migrate (regular Job, no hooks) - │ └── Uses openfga-migrator ServiceAccount + │ └── Runs as the OpenFGA ServiceAccount (or the migration one) │ └── Runs openfga migrate → succeeds ├── Creates ConfigMap with migrated version └── Pods pass readiness @@ -210,7 +210,7 @@ Nothing is deleted outright — every change is gated on `openfga-operator.enabl | `values.yaml`: `migration.serviceAccount.*` | Optional separate ServiceAccount for migration Jobs | | `values.yaml`: `migration.trigger` | Explicit rerun trigger for referenced Secret data changes | | `values.yaml`: migration pod values | `migrate.extraInitContainers`, `migrate.sidecars`, volumes, mounts, resources, timeout, non-hook annotations, and labels are forwarded to operator Jobs | -| `templates/serviceaccount.yaml`: second SA | Migration ServiceAccount | +| `templates/serviceaccount.yaml`: second SA | Optional migration ServiceAccount | | `charts/openfga-operator/` | Operator subchart (conditional dependency) | Users on `openfga-operator.enabled: false` (the default) see identical rendered output to the pre-operator chart, so gradual adoption is possible with no forced migration. @@ -219,9 +219,9 @@ Users on `openfga-operator.enabled: false` (the default) see identical rendered ### Positive -- **All 6 migration issues resolved** — no Helm hooks means no ArgoCD/FluxCD/`--wait` incompatibility +- **Migration issues resolved** (#211, #107, #120, #100, #126) — no Helm hooks means no ArgoCD/FluxCD/`--wait` incompatibility - **`k8s-wait-for` eliminated** — removes an unmaintained image with CVEs from the supply chain (#132, #144) -- **Least-privilege enforced** — separate ServiceAccounts for migration (DDL) and runtime (CRUD) (#95) +- **Least-privilege available** — `migration.serviceAccount.create` separates the migration (DDL) and runtime (CRUD) ServiceAccounts (#95) - **Runtime surface area reduced** — when `openfga-operator.enabled: true`, the legacy migration Job, init-container `k8s-wait-for` logic, and job-watching RBAC are skipped from the rendered manifest - **Migration is observable** — Job is a regular resource visible in all tools; ConfigMap records migration history; operator conditions surface errors - **Idempotent and crash-safe** — operator can restart at any point and resume correctly diff --git a/docs/adr/README.md b/docs/adr/README.md index 536a9d02..d6b3445e 100644 --- a/docs/adr/README.md +++ b/docs/adr/README.md @@ -10,7 +10,7 @@ We follow the format described by [Michael Nygard](https://cognitect.com/blog/20 | ADR | Title | Status | Date | |-----|-------|--------|------| -| [ADR-001](001-adopt-openfga-operator.md) | Adopt a Kubernetes Operator for OpenFGA Lifecycle Management | Accepted | 2026-04-06 | +| [ADR-001](001-adopt-openfga-operator.md) | Adopt a Kubernetes Operator for OpenFGA Lifecycle Management | Proposed | 2026-04-06 | | [ADR-002](002-operator-managed-migrations.md) | Replace Helm Hook Migrations with Operator-Managed Migrations | Proposed | 2026-04-06 | ---