Configuring Timeouts Behind Cloud Load Balancers
Gateways usually run behind a cloud load balancer, such as an AWS NLB or ALB, a Google Cloud load balancer, or an Azure Load Balancer or Application Gateway. Each hop on the path applies its own timeouts:
client ──▶ cloud load balancer ──▶ Envoy (TEG) ──▶ backend application
If those timeouts don't line up, you see intermittent problems that are hard to diagnose:
502errors from the cloud load balancer when it reuses a connection that Envoy has just closed.503errors from Envoy when it reuses a connection that the backend application has just closed.- Requests executed twice, or slow failures, because the load balancer and Envoy both retry.
- Long-lived connections (WebSockets, gRPC streams, idle keep-alive connections) dropped or reset with no error from Envoy.
504errors from Envoy or from the load balancer for requests that are slow but legitimate.- Connection errors during rollouts and scale-down.
This page lists the default timeouts in Envoy Gateway and in each major cloud load balancer, then recommends how to align them.
TEG uses the same timeout defaults as upstream Envoy Gateway. The TEG Helm charts don't override any of them, so everything on this page applies to both.
Envoy Gateway default timeouts
Some of these defaults are set explicitly by Envoy Gateway. Where Envoy Gateway leaves a setting unset, Envoy's own default applies.
Client side (cloud load balancer → Envoy)
Configure these in ClientTrafficPolicy.
| Timeout | Default | What it does | Field |
|---|---|---|---|
| HTTP connection idle timeout | 1 hour | Closes a downstream HTTP connection that has had no active requests for this long. | spec.timeout.http.idleTimeout |
| HTTP stream idle timeout | 5 minutes | Resets a request or stream with no activity in either direction for this long, which affects long polling, SSE and quiet WebSockets. When an HTTPRoute sets a request timeout, the route-level stream idle timeout is raised to max(1h, request timeout). | spec.timeout.http.streamIdleTimeout |
| Request received timeout | Disabled | Maximum time to receive the full request from the client. | spec.timeout.http.requestReceivedTimeout |
| Request headers timeout | Disabled | Maximum time to receive the request headers. | spec.timeout.http.requestHeadersReceivedTimeout |
| TCP/TLS listener idle timeout | 1 hour | Closes TCPRoute/TLSRoute connections with no bytes in either direction for this long. | spec.timeout.tcp.idleTimeout |
| Connection inspection timeout | 15 seconds | Maximum time for listener filters (TLS/SNI inspection, PROXY protocol). | spec.timeout.tcp.connectionInspectionTimeout |
| TLS handshake timeout | Not set (no limit) | Maximum time to complete the TLS handshake. | spec.timeout.tcp.tlsHandshakeTimeout |
| Max connection duration | Unlimited | Drains a connection after this age, whatever its activity. | spec.connection.connectionLimit.maxConnectionDuration |
| TCP keepalive | Off | When enabled without explicit values, Linux defaults apply: probes after 7200s idle, every 75s, 9 probes. | spec.tcpKeepalive (idleTime, interval, probes) |
| HTTP/2 PING keepalive | Off | Sends HTTP/2 PING frames on idle connections. | spec.http2.connectionKeepalive |
Route and backend side (Envoy → application)
Configure these on the HTTPRoute and in BackendTrafficPolicy.
| Timeout | Default | What it does | Field |
|---|---|---|---|
| Request (route) timeout | 15 seconds | Time from the end of the request until the full response arrives, retries included. Envoy Gateway leaves it unset, so Envoy's 15s default applies. When it expires the client gets a 504. | HTTPRoute rules[].timeouts.request / backendRequest, or BTP spec.timeout.http.requestTimeout |
| Upstream connect timeout | 10 seconds | Time to establish a TCP and TLS connection to a backend. | BTP spec.timeout.tcp.connectTimeout |
| Upstream connection idle timeout | 1 hour | Closes a pooled connection to a backend after it has been idle this long. This is longer than the keep-alive timeout of most application servers (see Case 4). | BTP spec.timeout.http.connectionIdleTimeout |
| Upstream max connection duration | Unlimited | BTP spec.timeout.http.maxConnectionDuration | |
| Max stream duration | Unlimited | BTP spec.timeout.http.maxStreamDuration | |
| Upstream TCP keepalive | Off | BTP spec.tcpKeepalive | |
| Retries | None | Envoy doesn't retry unless a retry policy is configured. A retry block without retryOn defaults to 2 retries on connect-failure, refused-stream, unavailable, cancelled and retriable-status-codes (503). See Retries across layers. | HTTPRoute rules[].retry, or BTP spec.retry |
The 15-second request timeout is the most common surprise. It is shorter than the default timeout of most cloud load balancers, so slow requests fail with a 504 from Envoy (response flag UT) before the load balancer gives up.
Shutdown and draining
Configure these in EnvoyProxy spec.shutdown.
| Setting | Default | What it does |
|---|---|---|
minDrainDuration | 10 seconds | Minimum time Envoy keeps serving while draining, even with no connections left. This gives load balancers time to notice that the pod is going away. |
drainTimeout | 60 seconds | Maximum time Envoy waits for connections to drain before exiting. |
healthCheckFailureDelay | 0 seconds | Delay before Envoy starts failing its readiness check during shutdown. |
Pod terminationGracePeriodSeconds | 360 seconds | Derived as drainTimeout + 300s. |
On shutdown, Envoy fails its readiness probe and starts draining. It sends Connection: close on HTTP/1.1 responses and GOAWAY on HTTP/2.
The Envoy Service is created with externalTrafficPolicy: Local by default, so cloud load balancers health-check each node's healthCheckNodePort and stop sending traffic to nodes with no ready Envoy pods.
Cloud load balancer default timeouts
These values come from the vendor documentation as of October 2026. Cloud providers change these settings from time to time, so check the linked pages before relying on a specific value.
AWS
| Network Load Balancer (NLB) | Application Load Balancer (ALB) | Classic ELB | |
|---|---|---|---|
| Layer | L4 (passthrough, no connection pooling) | L7 (reuses connections to targets) | L4/L7 |
| Idle timeout | 350s for TCP (configurable 60–6000s). TLS listeners: 350s, fixed. UDP: 120s, fixed. | 60s (1–4000s), applied on both the client and target side | 60s (1–4000s) |
| When the timeout expires | The flow is silently forgotten. The next packet from either side gets a TCP RST. | Connection closed. A target that closes an idle connection first can cause a 502. | Connection closed |
| Does TCP keepalive reset the timer? | Yes | No (HTTP/2 PING isn't supported either) | No |
| Request/response timeout | None | The idle timeout also bounds the wait for response bytes (504 when exceeded) | Same as idle |
| Other | Client keep-alive duration: 3600s, a maximum connection age | ||
| Target deregistration delay | 300s (0–3600s). Existing connections are not terminated afterwards unless deregistration_delay.connection_termination.enabled is true. | 300s | 300s (connection draining) |
| Health check defaults | 30s interval, 2 failures to mark unhealthy. The AWS Load Balancer Controller defaults to a 10s interval. | 30s interval, 5s timeout, 2 failures |
To configure these values with the AWS Load Balancer Controller:
- NLB idle timeout:
service.beta.kubernetes.io/aws-load-balancer-listener-attributes.TCP-443: tcp.idle_timeout.seconds=3600. To raise it above 350s, the target network interface's conntrackTcpEstablishedTimeoutmust be at least as large, or packets are dropped silently. - NLB target group attributes:
service.beta.kubernetes.io/aws-load-balancer-target-group-attributes: deregistration_delay.timeout_seconds=90,deregistration_delay.connection_termination.enabled=true - ALB:
alb.ingress.kubernetes.io/load-balancer-attributes: idle_timeout.timeout_seconds=120 - Classic ELB:
service.beta.kubernetes.io/aws-load-balancer-connection-idle-timeout: "120"
AWS recommends making the target's keep-alive timeout longer than the load balancer's idle timeout (ALB, troubleshooting 502s, NLB).
Google Cloud
External passthrough NLB (GKE LoadBalancer Service) | Internal passthrough NLB | Application Load Balancer (GKE Ingress / Gateway) | |
|---|---|---|---|
| Layer | L4 passthrough | L4 passthrough | L7 proxy (reuses connections to backends) |
| Idle timeout | 60s connection tracking, fixed | 600s connection tracking, configurable through connectionTrackingPolicy.idleTimeoutSec | Client keep-alive: 610s (5–1200s). Backend keep-alive: 600s, fixed. |
| When the timeout expires | No RST is sent. The next packet is re-hashed and normally reaches the same backend, unless the set of backends has changed. | Same as external | The proxy sends a FIN |
| Request/response timeout | None | None | Backend service timeout of 30s, which also bounds idle WebSockets (and all WebSockets on the classic ALB) |
| Retries | None | None | Global external: body-less requests (such as GET) that get a 502, 503 or 504 are retried once by default. Classic: GET requests that fail before response headers are retried once; not configurable. |
| Connection draining | Disabled (0s) by default, 0–3600s | Same | Disabled (0s) by default |
| Health check defaults (GKE) | Through healthCheckNodePort | NEG: 15s interval, 2 failures |
To configure these values on GKE:
- Ingress: BackendConfig
spec.timeoutSecandspec.connectionDraining.drainingTimeoutSec. - Gateway: GCPBackendPolicy
spec.default.timeoutSecandspec.default.connectionDraining.drainingTimeoutSec. - ALB client keep-alive: target proxy
--http-keep-alive-timeout-sec. - ALB retries (global external only): URL map
routeAction.retryPolicy(retryConditions,numRetries,perTryTimeout). A custom policy replaces the default one.
Google states that backends behind an Application Load Balancer must use an HTTP keep-alive timeout greater than 600 seconds, and gives 620s as an example (timeouts and retries, passthrough NLB).
Azure
Azure Load Balancer, Standard (AKS LoadBalancer Service) | Application Gateway v2 / AGIC | Application Gateway for Containers | Front Door | |
|---|---|---|---|---|
| Layer | L4 passthrough | L7 proxy | L7 proxy | L7 proxy (CDN) |
| Idle timeout | 4 min (4–100 min) | Frontend TCP: 4 min (4–30). Client HTTP/1.1 keep-alive: 120s, fixed. HTTP/2: 180s. Backend: not documented. | HTTP idle: 5 min, fixed | Client keep-alive: 90s, fixed |
| When the timeout expires | AKS enables TCP reset by default, so both sides receive an RST | Connection closed | Connection closed | Connection closed |
| Request/response timeout | None | 20s backend request timeout, which also applies to WebSockets | 60s request timeout, enforced even while data is streaming | 30s origin response timeout (16–240s) |
| Draining / probe failure | No draining. Established flows keep going to a backend whose probe fails until they close or go idle. | Connection draining off by default (1–3600s when enabled) | ||
| Health probe (AKS) | 5s interval, 2 failures | AGIC: 30s interval, 3 failures |
To configure these values:
- Azure LB idle timeout:
service.beta.kubernetes.io/azure-load-balancer-tcp-idle-timeout: "30"(in minutes). - AGIC request timeout:
appgw.ingress.kubernetes.io/request-timeout: "60"(seconds). - Application Gateway for Containers: HTTPRoute
timeouts.requestor RoutePolicyrouteTimeout.
Microsoft recommends TCP or application-level keepalives with an interval shorter than the idle timeout (TCP reset and idle timeout, AKS load balancer, Application Gateway settings, App Gateway for Containers).
How to align the timeouts
The recommended values and configuration depend on whether the load balancer is L4 passthrough (AWS NLB, Azure Load Balancer, GCP passthrough NLB) or an L7 proxy (AWS ALB, GCP Application Load Balancer, Azure Application Gateway).
Case 1: Behind an L7 load balancer, Envoy's idle timeout must be longer than the load balancer's
An L7 load balancer pools connections to Envoy and reuses them. If Envoy closes an idle connection at the moment the load balancer sends a new request on it, that request fails with a 502. The load balancer must always be the side that closes idle connections.
Envoy ClientTrafficPolicy timeout.http.idleTimeout should be greater than the load balancer's backend idle or keep-alive timeout, with some margin.
| Load balancer | LB backend idle timeout | Envoy idleTimeout |
|---|---|---|
| AWS ALB / Classic ELB | 60s (or your configured value) | Default 1h works. If you set it explicitly, use at least ALB value + 30s. |
| GCP Application Load Balancer | 600s, fixed | At least 620s. The default 1h works. |
| Azure App Gateway for Containers | 5 min | At least 330s. The default 1h works. |
| Azure Application Gateway | Not documented | Keep the default 1h. |
The Envoy default of 1 hour already meets this case. Problems usually start when someone lowers idleTimeout, for example to 300s, which is below the GCP backend keep-alive of 600s.
Case 2: Behind an L4 load balancer, keep connections active or close them before the load balancer forgets them
L4 load balancers don't pool connections, but they track every flow and drop it after a period of inactivity. The client and Envoy aren't told. The next packet on that connection is reset (AWS NLB, Azure LB) or, after a backend change, sent to the wrong node (GCP). Envoy's default idle timeouts of 1 hour are much longer than the NLB's 350s or Azure's 4 minutes, so idle keep-alive connections, WebSockets and gRPC streams are dropped silently.
Use one or both of these approaches:
- Make Envoy close idle connections first. Set the client
idleTimeout, andtcp.idleTimeoutfor TCP and TLS routes, below the load balancer's idle timeout. Envoy then closes cleanly with a FIN or GOAWAY, and clients reconnect without errors. - Keep long-lived connections alive. Enable
tcpKeepalivewith anidleTimewell below the load balancer's idle timeout. TCP keepalive resets the idle timer on the AWS NLB and on Azure LB. HTTP/2 PING keepalive (http2.connectionKeepalive) also works, because L4 load balancers see a PING as traffic.
You can also raise the load balancer's idle timeout, for example to 30 minutes on Azure or up to 6000s on an AWS NLB, and then set Envoy's values relative to the new number.
| Load balancer | LB idle timeout | Envoy client idleTimeout | Envoy tcpKeepalive.idleTime |
|---|---|---|---|
| AWS NLB (TCP) | 350s | 300s | 60s–120s |
| AWS NLB (TLS listener) | 350s, fixed | 300s | 60s–120s |
| Azure Load Balancer | 4 min (240s) | 200s | 60s–120s |
| GCP passthrough NLB | 60s (re-hashed, not reset) | Default is fine | Optional. Useful to survive node or backend changes. |
Case 3: Timeouts should get shorter the closer they are to the application
For request timeouts, each hop should allow at least as much time as the hop behind it. The innermost timeout then fires first and returns a clear error from Envoy, with the right response flags and access log entry. Otherwise the load balancer gives up while Envoy is still waiting.
client timeout ≥ cloud LB request timeout ≥ Envoy route timeout ≥ application timeout
Envoy's default route timeout is 15 seconds, which is lower than every cloud default. If any endpoint can take longer than 15 seconds, set the route timeout explicitly and make sure the load balancer allows more:
| Load balancer | LB request timeout | Action |
|---|---|---|
| AWS ALB | Idle timeout, 60s | Raise idle_timeout.timeout_seconds above the longest Envoy route timeout |
| GCP Application Load Balancer | 30s backend service timeout | Raise timeoutSec above the longest Envoy route timeout. For WebSockets, use a large value such as 3600s. |
| Azure Application Gateway | 20s | Raise request-timeout above the longest Envoy route timeout. This also applies to WebSockets. |
| Azure App Gateway for Containers | 60s, also applied to streaming | Raise it or set it to 0s for SSE or long downloads |
| Azure Front Door | 30s | Up to 240s |
| L4 load balancers | None | Only Case 2 applies |
For streaming, SSE and WebSocket routes, also check the stream idle timeout (5 minutes by default). Raise it with BackendTrafficPolicy timeout.http.streamIdleTimeout, or set a long route timeout, which raises the route's stream idle timeout with it.
Case 4: Envoy's upstream idle timeout must be shorter than the backend's keep-alive
Case 1 applies again on the next hop, with Envoy as the client. Envoy pools connections to each backend and keeps an idle one for 1 hour by default. Most application servers close idle keep-alive connections much sooner: Node.js after 5s, Go net/http servers when IdleTimeout is set, gunicorn after 2s, and many frameworks after 30–120s. When Envoy sends a request on a connection the backend is closing, the request fails with a 503 and response flag UC (upstream connection termination). These errors are intermittent and scale with traffic, so they're easy to miss.
Set BackendTrafficPolicy timeout.http.connectionIdleTimeout below the shortest keep-alive timeout of the backends it applies to. One policy for the whole Gateway only works if every backend has a longer keep-alive than the value you choose. For backends with very short keep-alives, either:
- attach a separate
BackendTrafficPolicyto their routes with a lower value, or - raise the keep-alive timeout in the application. This is usually better, because it avoids reconnecting on almost every request.
A retry on reset-before-request (see below) hides most of the remaining UC errors, but it doesn't prevent them. Fix the timeouts first.
Retries across layers
Retries and timeouts interact. Every retry adds another timeout's worth of latency, and when several layers retry, the attempts multiply. A cloud load balancer that retries twice in front of an Envoy that retries twice can send the same request to a backend up to nine times.
Retry in one layer, preferably Envoy: it knows whether any bytes reached the backend, so it can retry safely, and its access logs show every attempt. Limit the cloud load balancer to retrying connection-level failures, or nothing.
Envoy
Envoy Gateway doesn't configure any retries by default. For a retry that's safe for any method, including POST, use triggers that only fire when the request didn't reach the backend:
| Trigger | Fires when | Safe for non-idempotent requests? |
|---|---|---|
connect-failure | The upstream connection couldn't be established | Yes |
reset-before-request | The connection was reset before the request was sent, for example on a pooled connection the backend just closed | Yes |
refused-stream | The backend refused the HTTP/2 stream (REFUSED_STREAM) | Yes |
reset | The connection was reset at any point, including after the backend received the request | No |
5xx, gateway-error, retriable-status-codes | The backend or Envoy returned an error status | No. The backend may have processed the request. |
Adding a retry block without retryOn applies Envoy Gateway's default triggers. They include retriable-status-codes for 503, which replays requests the backend may already have processed. Always list the triggers explicitly.
Retries count against the route timeout, which covers all attempts. Use perRetry.timeout to bound each attempt.
Google Cloud Application Load Balancer
The global external Application Load Balancer retries by default: when no retry policy is configured, it retries body-less requests (such as GET) once on any 502, 503 or 504. That has two side effects:
- A request that times out waits for the backend service timeout twice. With the 30s default, a GET that hits a
504costs the client 60s. - A GET that reached the backend and failed afterwards runs twice. If it has side effects, such as firing webhooks, they happen twice.
To change this, set routeAction.retryPolicy on the URL map. A custom policy replaces the default one. Restrict retryConditions to failures where the request didn't reach Envoy, such as connect-failure and refused-stream, and let Envoy handle the rest.
In Google Cloud, numRetries is the total number of attempts, including the first one. numRetries: 2 means one retry, and numRetries: 1 disables retries. In Envoy, numRetries counts retries only, excluding the first attempt.
When no custom retry policy exists, the load balancer uses the backend service timeout for each attempt and ignores routeAction.timeout. The classic Application Load Balancer always retries GET requests that fail before response headers, as long as at least 80% of backends are healthy, and this can't be changed (timeouts and retries).
Draining during rollouts and scale-down
When an Envoy pod terminates, the load balancer needs time to notice before Envoy stops accepting connections:
minDrainDurationshould be at least as long as the load balancer needs to mark the target unhealthy (health check interval × unhealthy threshold), plus some margin. The 10s default is too short for most cloud load balancers. Examples: AWS NLB with default target group settings is 30s × 2 = 60s; GKE NEG is 15s × 2 = 30s; Azure LB is 5s × 2 = 10s.drainTimeoutshould cover your longest normal requests.terminationGracePeriodSecondsis adjusted todrainTimeout + 300sautomatically.- AWS NLB: the default deregistration delay of 300s doesn't close connections when it expires. Set
deregistration_delay.timeout_secondsto roughly matchdrainTimeout, and consider enablingderegistration_delay.connection_termination.enabled. - Azure Load Balancer never drains or resets established flows when a probe fails. Envoy closing connections during drain (with
Connection: closeorGOAWAY) is what moves clients to healthy pods. Make suredrainTimeoutis long enough for that.
Example configurations
The examples below apply to a Gateway named eg. Adjust names and namespaces to match your environment.
AWS NLB or Azure Load Balancer (L4)
apiVersion: gateway.envoyproxy.io/v1alpha1
kind: ClientTrafficPolicy
metadata:
name: cloud-lb-timeouts
namespace: default
spec:
targetRefs:
- group: gateway.networking.k8s.io
kind: Gateway
name: eg
timeout:
http:
idleTimeout: 300s # below NLB 350s; use 200s for Azure LB (240s)
tcp:
idleTimeout: 300s # for TCPRoute/TLSRoute listeners
tcpKeepalive:
idleTime: 60s # keep long-lived connections alive at L4
interval: 30s
probes: 3
AWS ALB or GCP Application Load Balancer (L7)
apiVersion: gateway.envoyproxy.io/v1alpha1
kind: ClientTrafficPolicy
metadata:
name: cloud-lb-timeouts
namespace: default
spec:
targetRefs:
- group: gateway.networking.k8s.io
kind: Gateway
name: eg
timeout:
http:
idleTimeout: 3600s # must stay above the LB backend keep-alive (ALB 60s, GCP 600s)
Request timeouts
Set the route timeout to what the application needs, and raise the load balancer's timeout above it:
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
name: backend
spec:
parentRefs:
- name: eg
rules:
- backendRefs:
- name: backend
port: 3000
timeouts:
request: 50s # Envoy default is 15s; keep below the LB timeout (e.g. ALB 60s)
Upstream idle timeout and safe retries
Apply a Gateway-wide default, then override it for backends with very short keep-alives:
apiVersion: gateway.envoyproxy.io/v1alpha1
kind: BackendTrafficPolicy
metadata:
name: upstream-defaults
namespace: default
spec:
targetRefs:
- group: gateway.networking.k8s.io
kind: Gateway
name: eg
timeout:
http:
connectionIdleTimeout: 25s # below the shortest backend keep-alive covered by this policy
retry:
numRetries: 1 # Envoy counts retries, excluding the first attempt
retryOn:
triggers: # only fire when the request didn't reach the backend
- connect-failure
- reset-before-request
- refused-stream
---
apiVersion: gateway.envoyproxy.io/v1alpha1
kind: BackendTrafficPolicy
metadata:
name: short-keepalive-backend
namespace: default
spec:
targetRefs:
- group: gateway.networking.k8s.io
kind: HTTPRoute
name: node-app # backend closes idle connections after 5s
mergeType: StrategicMerge # keep the Gateway-level retry policy
timeout:
http:
connectionIdleTimeout: 4s
Without mergeType, a route-level BackendTrafficPolicy replaces the Gateway-level one for that route, and the route loses the retry policy.
Google Cloud ALB retry policy
Limit the global external Application Load Balancer to connection-level retries in the URL map, and leave the rest to Envoy:
defaultRouteAction:
retryPolicy:
retryConditions:
- connect-failure
- refused-stream
numRetries: 2 # total attempts in Google Cloud: one retry
Load balancer settings and graceful shutdown
Set load balancer annotations and shutdown timing on the EnvoyProxy resource referenced by your GatewayClass. When you install TEG with Helm, set them under config.envoyProxy in the teg-envoy-gateway-helm values. The example below is for an AWS NLB:
apiVersion: gateway.envoyproxy.io/v1alpha1
kind: EnvoyProxy
metadata:
name: teg-envoy-proxy
namespace: envoy-gateway-system
spec:
shutdown:
minDrainDuration: 60s # >= LB health check interval x unhealthy threshold
drainTimeout: 90s
provider:
type: Kubernetes
kubernetes:
envoyService:
annotations:
service.beta.kubernetes.io/aws-load-balancer-type: external
service.beta.kubernetes.io/aws-load-balancer-nlb-target-type: ip
service.beta.kubernetes.io/aws-load-balancer-target-group-attributes: >-
deregistration_delay.timeout_seconds=90,deregistration_delay.connection_termination.enabled=true
# For Azure instead:
# service.beta.kubernetes.io/azure-load-balancer-tcp-idle-timeout: "30"
Troubleshooting with access logs
Envoy access logs show which hop closed the connection:
| Symptom | Likely cause |
|---|---|
502 at the cloud load balancer, with no matching Envoy access log entry | Envoy closed an idle connection that the load balancer reused. Raise Envoy's idleTimeout (Case 1). |
504 from Envoy with response flag UT | The route timeout expired. The default is 15s (Case 3). |
504 at the load balancer while the Envoy log shows a long %DURATION% or no entry yet | The load balancer's request timeout is shorter than Envoy's route timeout (Case 3). |
Clients see connection reset on reused or long-lived connections | An L4 load balancer idle timeout expired (Case 2). |
Response flag SI | The stream idle timeout expired (5 min by default). |
503 with response flag UC | The backend application closed a pooled connection that Envoy reused. Set BackendTrafficPolicy timeout.http.connectionIdleTimeout below the application's keep-alive timeout (Case 4). |
| GET requests reach the backend twice, or side effects (such as webhooks) fire twice | The Google Cloud global external ALB default retry replayed a request that got a 502, 503 or 504. Set a custom retryPolicy (Retries across layers). |
A 504 takes twice the configured timeout (for example 60s with a 30s timeout) | A retry layer waited for the full timeout on each attempt, such as the Google Cloud ALB default retry. Restrict retries to connection failures, or set perRetry.timeout in Envoy. |
Several Envoy access log entries for one client request, or %UPSTREAM_REQUEST_ATTEMPT_COUNT% above 1 | Retries in more than one layer multiplied. Keep retries in Envoy and limit the load balancer to connection-level failures. |