Managing Gateways during Upgrades
When you upgrade your TSB ControlPlane, the new operator may regenerate gateway manifests with updated values — a new proxy image, new environment variables, a new injection template. Applying those changes restarts the gateway pods, and during a large upgrade many gateways restarting at once can mean unplanned downtime.
To prevent this, an upgrade does not restart your gateways. The new operator checks each gateway, and if the upgrade would change it, the gateway is marked RECONCILIATION_DIRTY and keeps running on its current configuration. You then decide when each gateway picks up the change.
Platform Operators have two tools to manage this:
| Tool | What it does |
|---|---|
install.tetrate.io/reconcile-before label on a gateways.install.tetrate.io CR | Tells the operator to apply pending changes to that gateway. Use it to release gateways in batches after an upgrade, or — set before an upgrade — to have gateways update automatically. |
/debug/gateway-reconcile-diff endpoint on xcp-operator-edge | Dry-run preview of exactly what would change for each gateway if it were reconciled. |
A gateway is reconciled when the operator (re)applies its generated manifests. A reconcile only restarts the gateway's pods when the new manifests change the pod template — for example, a different proxy image or new environment variables.
How gateways behave during an upgrade
After the new operator starts, every gateway ends up in one of these states:
| The upgrade... | Gateway phase | What happens to the pods |
|---|---|---|
| Does not change the gateway's manifests | READY | Nothing. The gateway is already up to date. |
| Changes the gateway's manifests | RECONCILIATION_DIRTY | Nothing yet. The gateway keeps running on its current configuration until you release it. |
Changes the gateway's manifests, and the gateway has a reconcile-before label with a date in the future | READY | The changes are applied right away, and the pods restart if the pod template changed. This is the behavior of TSB releases before this feature. |
You choose between two workflows:
- Release gateways gradually (recommended, and the default). Upgrade, see which gateways are dirty, preview the changes, then release them in batches during your maintenance windows.
- Update gateways automatically. Before the upgrade, set the
reconcile-beforelabel to a date after your upgrade window. Gateways are updated as soon as the new operator starts, as they were in earlier releases.
This behavior applies to gateways installed with the unified gateways.install.tetrate.io resource. Gateways installed with the legacy IngressGateway, Tier1Gateway and EgressGateway resources are always reconciled on upgrade, and restart if their pod template changed. To control when those restart, upgrade them to the unified Gateway first.
A dirty gateway only holds back changes that come from the upgrade. If you edit the gateway's gateways.install.tetrate.io CR spec, the operator reconciles it, and that reconcile applies the pending upgrade changes as well. Plan spec edits to dirty gateways as if they were a release.
The first upgrade from these versions needs one extra step before you start. See Upgrading from TSB 1.12.12 or earlier.
Recommended Upgrade Workflow
Follow these steps for every cluster where you want to control when gateways restart.
Upgrade the ControlPlane
If this is your first upgrade from TSB 1.12.12 or earlier, scale down the old edge operator first, as described in Upgrading from TSB 1.12.12 or earlier.
Follow your normal TSB upgrade procedure. You don't need to do anything to your gateways beforehand. When the new operator starts, gateways the upgrade would change move to
RECONCILIATION_DIRTYand keep running on their current configuration.Find the dirty gateways
List the phase of every gateway. The gateway status lives on the
GatewayDeploymentresource the operator creates for eachgateways.install.tetrate.ioCR, in the ControlPlane namespace. The command below shows the namespace and name of thegateways.install.tetrate.ioCR it belongs to:List gateway phaseskubectl get gatewaydeployments -A \
-o custom-columns='NAMESPACE:.metadata.labels.install\.tetrate\.io/owner-namespace,GATEWAY:.metadata.labels.install\.tetrate\.io/owner-name,PHASE:.status.phase'
# NAMESPACE GATEWAY PHASE
# bookinfo bookinfo-gw RECONCILIATION_DIRTY
# tier1 tier1-gw RECONCILIATION_DIRTY
# envoy-staging staging-gw READYREADYgateways need no action. EveryRECONCILIATION_DIRTYgateway has changes waiting to be applied.Preview what the upgrade would change
Before releasing any gateway, check why it is dirty and what reconciling it would change.
Check the gateway status. For a quick look at one gateway, read its
GatewayDeploymentstatus. Find it by the name and namespace of itsgateways.install.tetrate.ioCR:Show why a gateway is dirtykubectl get gatewaydeployments -n istio-system \
-l install.tetrate.io/owner-name=bookinfo-gw,install.tetrate.io/owner-namespace=bookinfo -o yamlThe
DirtyStateDetectedcondition instatusreports that the gateway's generated configuration differs from what is applied, and lists the gateway's resources:status:
phase: RECONCILIATION_DIRTY
conditions:
- type: DirtyStateDetected
status: "True"
reason: ContentHashMismatch
message: 'Dirty resources: Deployment/bookinfo/bookinfo-gw, Service/bookinfo/bookinfo-gw, ...'Preview the exact changes. The status tells you a gateway is dirty, but not what changed or whether its pods would restart. For that, query the dry-run diff endpoint. The endpoint is read-only — it never modifies cluster state.
Port-forward to the edge operator:
Port-forward to xcp-operator-edgekubectl port-forward -n istio-system deployment/xcp-operator-edge 8090:8090Then query the diff endpoint:
http://localhost:8090/debug/gateway-reconcile-diffTo narrow the diff, pass
namespacefor all gateways in a namespace, ornamespaceandnamefor a single gateway (namerequiresnamespace):http://localhost:8090/debug/gateway-reconcile-diff?namespace=bookinfo&name=bookinfo-gwThese filters match the gateway's Kubernetes Deployment, which has the same name and namespace as its
gateways.install.tetrate.ioCR. In the response,deploymentName/deploymentNamespaceidentify that Deployment.name/namespaceidentify the operator'sGatewayDeploymentresource, which is named after the gateway type and UID and lives in the ControlPlane namespace:{
"gateways": [
{
"name": "unified-6f1c2a9e-3b7d-4c1e-9a52-0d8e4f7b1c23",
"namespace": "istio-system",
"deploymentName": "bookinfo-gw",
"deploymentNamespace": "bookinfo",
"kind": "GatewayDeployment",
"revision": "default",
"reconcileEnabled": true,
"diff": {
"deployment": {
"hasChanges": true,
"summary": "pod template changed (will cause restart): ...",
"willCauseRestart": true
},
"service": { "hasChanges": false },
"serviceAccount": { "hasChanges": false },
"hpa": { "hasChanges": false }
}
}
],
"summary": {
"total": 47,
"withChanges": 3,
"willCauseRestart": 2
}
}Filtering by name when in-place gateway upgrade is disabledWhen in-place gateway upgrade is disabled, the Deployment name carries the revision suffix (for example
my-gateway-stablerather thanmy-gateway). Use the exactdeploymentNamevalue — visible in an unfiltered response — when filtering byname.Use
jqto quickly list the gateways that would restart:List gateways whose pods would restart on reconcilecurl -s "http://localhost:8090/debug/gateway-reconcile-diff" \
| jq -r '.gateways[]
| select([.diff // {} | .[]? | select(. != null) | .willCauseRestart == true] | any)
| "\(.deploymentNamespace)/\(.deploymentName)"'
# bookinfo/bookinfo-gw
# tier1/tier1-gwOr get a per-resource summary of the changes:
Per-resource change summarycurl -s "http://localhost:8090/debug/gateway-reconcile-diff" \
| jq -r '.gateways[] | . as $g | (.diff // {}) | to_entries[]
| select(.value != null and .value.hasChanges == true)
| "\($g.deploymentNamespace)/\($g.deploymentName)\t\(.key)\t\(.value.summary // "")"'Slow on large clusters?On clusters with many gateways the cluster-wide diff can take time. Increase the worker concurrency:
http://localhost:8090/debug/gateway-reconcile-diff?concurrency=5The endpoint is for debugging only — for very large responses, consider raising the memory limit on the
xcp-operator-edgepod before requesting a cluster-wide diff.Release gateways during your maintenance window
When you are ready for gateways to update, release them in batches so you can validate each batch before moving on.
You release a gateway by setting the
install.tetrate.io/reconcile-beforelabel on itsgateways.install.tetrate.ioCR. As the name says, the label tells the operator to reconcile the gateway before a deadline. When you set it, the operator applies the pending changes right away, and the pods restart if the pod template changed. After the deadline passes, the label has no effect.The deadline only needs to be later than now. Setting it one hour from now leaves plenty of time. The value is a UTC timestamp in the format
YYYYMMDDThhmmssZ— for example,20261001T120000Zmeans 1 October 2026, 12:00 UTC. Generate one for an hour from now with:A timestamp one hour from now# Linux (GNU date)
DEADLINE=$(date -u -d '+1 hour' +%Y%m%dT%H%M%SZ)
# macOS (BSD date)
DEADLINE=$(date -u -v+1H +%Y%m%dT%H%M%SZ)The label is already thereTSB adds the
reconcile-beforelabel to everygateways.install.tetrate.ioCR, set to the release date of your TSB version. A date in the past does nothing, so it is safe to leave. Always pass--overwritewhen you set it.Release a single gateway:
kubectl label gateways.install.tetrate.io bookinfo-gw -n bookinfo \
install.tetrate.io/reconcile-before=$DEADLINE --overwriteRelease every gateway in a namespace:
kubectl label gateways.install.tetrate.io --all -n envoy-prod-us-east \
install.tetrate.io/reconcile-before=$DEADLINE --overwriteRelease everything:
kubectl label gateways.install.tetrate.io --all -A \
install.tetrate.io/reconcile-before=$DEADLINE --overwriteReleasing a gateway that is already
READYis harmless. The operator re-applies the same manifests and the pods don't restart.If you manage
gateways.install.tetrate.ioCRs with GitOps, set the label in your manifests instead, so your GitOps tool doesn't remove it. The label is safe to leave in place after the timestamp passes.Watch the gateway phases transition from
RECONCILIATION_DIRTYtoREADYas each batch is released:kubectl get gatewaydeployments -A \
-o custom-columns='NAMESPACE:.metadata.labels.install\.tetrate\.io/owner-namespace,GATEWAY:.metadata.labels.install\.tetrate\.io/owner-name,PHASE:.status.phase'If a released gateway stays in
RECONCILIATION_DIRTY, see Troubleshooting: a gateway stays in RECONCILIATION_DIRTY.Confirm gateways are healthy and on the new proxy version
After release, check the rollout completed and all pods are running. In the
kubectl get deploymentsoutput, theREADYcolumn should shown/n(every replica ready); inkubectl get pods, every pod should beRunningand1/1.kubectl get deployments -n bookinfo
kubectl get pods -n bookinfoThen confirm the
istio-proxycontainer is running the image that ships with the new ControlPlane. Each gateway's deployment carries anapp: tsb-gateway-...label whose exact value depends on the gateway and namespace — discover it withkubectl get deploy -n <namespace> --show-labels, then use it to select the pods:Show the running proxy image for a gatewaykubectl get pods -n bookinfo -l app=tsb-gateway-bookinfo \
-o jsonpath='{range .items[*]}{.metadata.name}{"\t"}{.spec.containers[?(@.name=="istio-proxy")].image}{"\n"}{end}'
# bookinfo-gw-7c8b9d5f4-abc12 docker.io/istio/proxyv2:1.24.4
# bookinfo-gw-7c8b9d5f4-def34 docker.io/istio/proxyv2:1.24.4Gateways with no pod-template change won't restart on their ownIf a TSB upgrade does not change a gateway's pod template — for example, the proxy image tag is unchanged — the gateway stays
READYand its pods keep running the previous image until you restart them yourself.To pick up a new image on those gateways, do a rolling restart of the underlying Kubernetes Deployment during your maintenance window:
Force-restart a gateway's podskubectl rollout restart deployment <gateway-deployment-name> -n <namespace>
Updating gateways automatically during an upgrade
If you prefer the behavior of earlier TSB releases — every gateway picks up the upgrade's changes as soon as the new operator starts — set the reconcile-before label before you upgrade, to a time after your upgrade will finish.
For example, if you plan to upgrade on 1 October 2026, a deadline at the end of that day covers the whole window:
kubectl label gateways.install.tetrate.io --all -A \
install.tetrate.io/reconcile-before=20261001T235959Z --overwrite
When the new operator starts before the deadline, it applies each gateway's changes immediately, and gateways whose pod template changed restart. No gateway goes to RECONCILIATION_DIRTY. Once the deadline passes, the label has no effect, so your next upgrade goes back to the default behavior unless you set a new deadline.
You can mix both workflows. For example, label only a staging namespace so it updates automatically, then release production namespaces gradually.
The label only has an effect while its deadline is in the future. If the deadline passes before the new operator is running, the label has no effect and changed gateways go to RECONCILIATION_DIRTY. You can still release them afterwards with the recommended workflow.
Setting the label also makes the current operator reconcile those gateways right away. Any gateway still RECONCILIATION_DIRTY from a previous upgrade is released at that moment, so check for dirty gateways before you set it.
A deadline far in the future keeps gateways updating automatically on every upgrade. It also means every reconcile of those gateways re-applies their full manifests, as in earlier releases.
Upgrading from TSB 1.12.12 or earlier
This only applies the first time you upgrade from a TSB version earlier than 1.12.13 to 1.12.13 or later.
During an upgrade, the old xcp-operator-edge keeps running until the new one replaces it. The old operator does not have RECONCILIATION_DIRTY support. If the upgrade updates the gateway resources before the old operator is replaced, the old operator can reconcile them and restart every gateway — the outcome you are trying to avoid.
To remove this race, scale the old operator down just before you upgrade:
kubectl scale deployment xcp-operator-edge -n istio-system --replicas=0
kubectl get deployment xcp-operator-edge -n istio-system
# NAME READY UP-TO-DATE AVAILABLE AGE
# xcp-operator-edge 0/0 0 0 200d
Your gateways keep serving traffic while it is scaled down. Only operator management of gateways stops.
Then upgrade the ControlPlane. The upgrade deploys the new xcp-operator-edge and scales it back up for you. Confirm the new operator is running before you continue:
kubectl get deployment xcp-operator-edge -n istio-system \
-o custom-columns='READY:.status.readyReplicas,IMAGE:.spec.template.spec.containers[0].image'
If READY still shows 0 or <none> once the upgrade has finished, scale it back up yourself:
kubectl scale deployment xcp-operator-edge -n istio-system --replicas=1
Later upgrades do not need this step, because the running operator already supports RECONCILIATION_DIRTY.
Troubleshooting: a gateway stays in RECONCILIATION_DIRTY
RECONCILIATION_DIRTY means the operator detected changes for the gateway that have not been applied. The gateway keeps running on its existing configuration in the meantime. This is expected after an upgrade, until you release the gateway.
If you set the reconcile-before label and the gateway is still dirty, check the following:
-
The label is on the right resource. Set it on the
gateways.install.tetrate.ioCR in the gateway's namespace, not on theGatewayDeploymentor the Kubernetes Deployment. TSB copies it to theGatewayDeploymentfor you. -
The timestamp is in the future, in UTC. The value must match
YYYYMMDDThhmmssZexactly, for example20261001T120000Z. A timestamp in the past, or one written in local time that is already past in UTC, has no effect. The operator logs aFailed to parsewarning for a malformed value. -
The reconcile succeeded. If applying the changes fails, the gateway stays dirty and the operator retries. Check the operator logs:
kubectl logs -n istio-system deployment/xcp-operator-edge | grep <gateway-name>
To see why the gateway is dirty and what would change, check its status and the diff endpoint as described in Preview what the upgrade would change.
Reference
Gateway phases
GatewayDeployment.status.phase reports the overall state of a gateway:
| Phase | Meaning |
|---|---|
READY | The gateway is reconciled and the workload is healthy. |
PENDING | The gateway is being created or its workload is still coming up. |
WAITING_FOR_LOAD_BALANCER | The service has been created and is waiting for the cloud load balancer to be assigned. |
RECONCILIATION_DIRTY | The operator detected changes for the gateway that have not been applied, typically after an upgrade. The workload keeps running on its existing configuration. Release it with the reconcile-before label. |
TRANSLATION_FAILED | TSB could not translate the gateway's configuration into Istio resources. |
The reconcile-before label
| Label | install.tetrate.io/reconcile-before |
| Set on | The gateways.install.tetrate.io CR |
| Format | UTC timestamp, YYYYMMDDThhmmssZ (for example 20261001T120000Z) |
| Before the timestamp | Every reconcile of the gateway applies its full manifests. Pending changes are applied, and the pods restart if the pod template changed. |
| After the timestamp | No effect. Safe to leave in place, including in GitOps-managed manifests. |
| Default | TSB sets it to the release date of your TSB version when it is missing. TSB never overwrites a value you set. |
Observability
The edge operator exposes Prometheus metrics for gateway reconciliation. Scrape them from the operator service or port-forward locally:
kubectl port-forward -n istio-system service/xcp-operator-edge 8084:8080
curl -s http://localhost:8084/metrics | grep -E 'gateway_(reconcile|dirty|force)'
All gateway metrics carry the labels gateway_type, gateway_namespace, and gateway_name. The most useful ones for upgrade workflows:
| Metric | Type | Description |
|---|---|---|
gateway_dirty_state | Gauge | 1 when a gateway is RECONCILIATION_DIRTY, 0 otherwise. |
gateway_reconcile_skipped_total | Counter | Incremented each time the operator skips reconciling a gateway. The reason label is already_reconciled (nothing to do) or dirty_skipped (changes are waiting for release). |
gateway_force_reconcile_total | Counter | Incremented each time the install.tetrate.io/reconcile-before label triggers a reconcile. |
Example alert rules:
| Alert | PromQL | For | Severity |
|---|---|---|---|
GatewayDirtyTooLong | gateway_dirty_state == 1 | 7d | Warning — a gateway was likely never released after an upgrade. |
Grafana Dashboard
As an alternative to querying these metrics directly, you can use the XCP Edge Operator Grafana dashboard (uid xcp-edge-operator). Its Gateway Upgrades row is built specifically for this workflow, so you can watch a rollout progress without leaving Grafana.
The process to install and operate TSB internal dashboards is explained in the Telemetry documentation.
Two template variables scope every panel on the dashboard:
cluster— the XCP Edge cluster to inspect.namespace— the gateway namespace to inspect (supports the panels that break results out per-gateway).
The Gateway Upgrades row surfaces:
| Panel | What it tells you |
|---|---|
| Installed Gateways | Gateways actively managed by the current Istio revision, by gateway type. |
| Gateways Ignored (Revision Mismatch) | Gateways belonging to a different Istio revision. Non-zero means an upgrade is in progress. |
| Gateways in Dirty State | Gateways with changes waiting to be released. Non-zero after an upgrade is expected, and should fall to zero as you release batches. See Troubleshooting: a gateway stays in RECONCILIATION_DIRTY. |
| Gateway Reconcile Rate (Success vs Failure) | Reconcile throughput by gateway type. Failures during an upgrade can indicate incompatible configuration or image-pull issues. |
| Gateway Reconcile Time (p50/p95/p99) | Reconcile latency percentiles. Elevated latency during an upgrade can indicate resource contention or slow image pulls. |
| Skipped Reconciliations by Reason | Skip rate broken out by reason: already_reconciled (nothing to do) and dirty_skipped (changes waiting for release). already_reconciled is expected to dominate at steady state. |
| Force Reconcile Triggers | Rate of reconciles triggered by the install.tetrate.io/reconcile-before label — each one is a gateway you released. |
| Time Since Last Gateway Reconcile | Age of the last reconcile per gateway. A large value can mean a gateway is stuck mid-upgrade. |
| Gateways in Dirty State (table) | Per-gateway list (filtered by namespace) so you can identify exactly which gateways need attention, not just the aggregate count. |
A practical way to use the dashboard during a rollout: set namespace to the batch you just released, watch Gateways in Dirty State fall to zero and Force Reconcile Triggers spike as the batch picks up the changes, and confirm Gateway Reconcile Rate shows no failures.