Skip to main content
logoTetrate Enterprise Gateway for EnvoyVersion: v1.9.x

Configuring Timeouts Behind Cloud Load Balancers

Gateways usually run behind a cloud load balancer, such as an AWS NLB or ALB, a Google Cloud load balancer, or an Azure Load Balancer or Application Gateway. Each hop on the path applies its own timeouts:

client ──▶ cloud load balancer ──▶ Envoy (TEG) ──▶ backend application

If those timeouts don't line up, you see intermittent problems that are hard to diagnose:

  • 502 errors from the cloud load balancer when it reuses a connection that Envoy has just closed.
  • 503 errors from Envoy when it reuses a connection that the backend application has just closed.
  • Requests executed twice, or slow failures, because the load balancer and Envoy both retry.
  • Long-lived connections (WebSockets, gRPC streams, idle keep-alive connections) dropped or reset with no error from Envoy.
  • 504 errors from Envoy or from the load balancer for requests that are slow but legitimate.
  • Connection errors during rollouts and scale-down.

This page lists the default timeouts in Envoy Gateway and in each major cloud load balancer, then recommends how to align them.

info

TEG uses the same timeout defaults as upstream Envoy Gateway. The TEG Helm charts don't override any of them, so everything on this page applies to both.

Envoy Gateway default timeouts​

Some of these defaults are set explicitly by Envoy Gateway. Where Envoy Gateway leaves a setting unset, Envoy's own default applies.

Client side (cloud load balancer → Envoy)​

Configure these in ClientTrafficPolicy.

TimeoutDefaultWhat it doesField
HTTP connection idle timeout1 hourCloses a downstream HTTP connection that has had no active requests for this long.spec.timeout.http.idleTimeout
HTTP stream idle timeout5 minutesResets a request or stream with no activity in either direction for this long, which affects long polling, SSE and quiet WebSockets. When an HTTPRoute sets a request timeout, the route-level stream idle timeout is raised to max(1h, request timeout).spec.timeout.http.streamIdleTimeout
Request received timeoutDisabledMaximum time to receive the full request from the client.spec.timeout.http.requestReceivedTimeout
Request headers timeoutDisabledMaximum time to receive the request headers.spec.timeout.http.requestHeadersReceivedTimeout
TCP/TLS listener idle timeout1 hourCloses TCPRoute/TLSRoute connections with no bytes in either direction for this long.spec.timeout.tcp.idleTimeout
Connection inspection timeout15 secondsMaximum time for listener filters (TLS/SNI inspection, PROXY protocol).spec.timeout.tcp.connectionInspectionTimeout
TLS handshake timeoutNot set (no limit)Maximum time to complete the TLS handshake.spec.timeout.tcp.tlsHandshakeTimeout
Max connection durationUnlimitedDrains a connection after this age, whatever its activity.spec.connection.connectionLimit.maxConnectionDuration
TCP keepaliveOffWhen enabled without explicit values, Linux defaults apply: probes after 7200s idle, every 75s, 9 probes.spec.tcpKeepalive (idleTime, interval, probes)
HTTP/2 PING keepaliveOffSends HTTP/2 PING frames on idle connections.spec.http2.connectionKeepalive

Route and backend side (Envoy → application)​

Configure these on the HTTPRoute and in BackendTrafficPolicy.

TimeoutDefaultWhat it doesField
Request (route) timeout15 secondsTime from the end of the request until the full response arrives, retries included. Envoy Gateway leaves it unset, so Envoy's 15s default applies. When it expires the client gets a 504.HTTPRoute rules[].timeouts.request / backendRequest, or BTP spec.timeout.http.requestTimeout
Upstream connect timeout10 secondsTime to establish a TCP and TLS connection to a backend.BTP spec.timeout.tcp.connectTimeout
Upstream connection idle timeout1 hourCloses a pooled connection to a backend after it has been idle this long. This is longer than the keep-alive timeout of most application servers (see Case 4).BTP spec.timeout.http.connectionIdleTimeout
Upstream max connection durationUnlimitedBTP spec.timeout.http.maxConnectionDuration
Max stream durationUnlimitedBTP spec.timeout.http.maxStreamDuration
Upstream TCP keepaliveOffBTP spec.tcpKeepalive
RetriesNoneEnvoy doesn't retry unless a retry policy is configured. A retry block without retryOn defaults to 2 retries on connect-failure, refused-stream, unavailable, cancelled and retriable-status-codes (503). See Retries across layers.HTTPRoute rules[].retry, or BTP spec.retry
warning

The 15-second request timeout is the most common surprise. It is shorter than the default timeout of most cloud load balancers, so slow requests fail with a 504 from Envoy (response flag UT) before the load balancer gives up.

Shutdown and draining​

Configure these in EnvoyProxy spec.shutdown.

SettingDefaultWhat it does
minDrainDuration10 secondsMinimum time Envoy keeps serving while draining, even with no connections left. This gives load balancers time to notice that the pod is going away.
drainTimeout60 secondsMaximum time Envoy waits for connections to drain before exiting.
healthCheckFailureDelay0 secondsDelay before Envoy starts failing its readiness check during shutdown.
Pod terminationGracePeriodSeconds360 secondsDerived as drainTimeout + 300s.

On shutdown, Envoy fails its readiness probe and starts draining. It sends Connection: close on HTTP/1.1 responses and GOAWAY on HTTP/2.

The Envoy Service is created with externalTrafficPolicy: Local by default, so cloud load balancers health-check each node's healthCheckNodePort and stop sending traffic to nodes with no ready Envoy pods.

Cloud load balancer default timeouts​

These values come from the vendor documentation as of October 2026. Cloud providers change these settings from time to time, so check the linked pages before relying on a specific value.

AWS​

Network Load Balancer (NLB)Application Load Balancer (ALB)Classic ELB
LayerL4 (passthrough, no connection pooling)L7 (reuses connections to targets)L4/L7
Idle timeout350s for TCP (configurable 60–6000s). TLS listeners: 350s, fixed. UDP: 120s, fixed.60s (1–4000s), applied on both the client and target side60s (1–4000s)
When the timeout expiresThe flow is silently forgotten. The next packet from either side gets a TCP RST.Connection closed. A target that closes an idle connection first can cause a 502.Connection closed
Does TCP keepalive reset the timer?YesNo (HTTP/2 PING isn't supported either)No
Request/response timeoutNoneThe idle timeout also bounds the wait for response bytes (504 when exceeded)Same as idle
OtherClient keep-alive duration: 3600s, a maximum connection age
Target deregistration delay300s (0–3600s). Existing connections are not terminated afterwards unless deregistration_delay.connection_termination.enabled is true.300s300s (connection draining)
Health check defaults30s interval, 2 failures to mark unhealthy. The AWS Load Balancer Controller defaults to a 10s interval.30s interval, 5s timeout, 2 failures

To configure these values with the AWS Load Balancer Controller:

  • NLB idle timeout: service.beta.kubernetes.io/aws-load-balancer-listener-attributes.TCP-443: tcp.idle_timeout.seconds=3600. To raise it above 350s, the target network interface's conntrack TcpEstablishedTimeout must be at least as large, or packets are dropped silently.
  • NLB target group attributes: service.beta.kubernetes.io/aws-load-balancer-target-group-attributes: deregistration_delay.timeout_seconds=90,deregistration_delay.connection_termination.enabled=true
  • ALB: alb.ingress.kubernetes.io/load-balancer-attributes: idle_timeout.timeout_seconds=120
  • Classic ELB: service.beta.kubernetes.io/aws-load-balancer-connection-idle-timeout: "120"

AWS recommends making the target's keep-alive timeout longer than the load balancer's idle timeout (ALB, troubleshooting 502s, NLB).

Google Cloud​

External passthrough NLB (GKE LoadBalancer Service)Internal passthrough NLBApplication Load Balancer (GKE Ingress / Gateway)
LayerL4 passthroughL4 passthroughL7 proxy (reuses connections to backends)
Idle timeout60s connection tracking, fixed600s connection tracking, configurable through connectionTrackingPolicy.idleTimeoutSecClient keep-alive: 610s (5–1200s). Backend keep-alive: 600s, fixed.
When the timeout expiresNo RST is sent. The next packet is re-hashed and normally reaches the same backend, unless the set of backends has changed.Same as externalThe proxy sends a FIN
Request/response timeoutNoneNoneBackend service timeout of 30s, which also bounds idle WebSockets (and all WebSockets on the classic ALB)
RetriesNoneNoneGlobal external: body-less requests (such as GET) that get a 502, 503 or 504 are retried once by default. Classic: GET requests that fail before response headers are retried once; not configurable.
Connection drainingDisabled (0s) by default, 0–3600sSameDisabled (0s) by default
Health check defaults (GKE)Through healthCheckNodePortNEG: 15s interval, 2 failures

To configure these values on GKE:

  • Ingress: BackendConfig spec.timeoutSec and spec.connectionDraining.drainingTimeoutSec.
  • Gateway: GCPBackendPolicy spec.default.timeoutSec and spec.default.connectionDraining.drainingTimeoutSec.
  • ALB client keep-alive: target proxy --http-keep-alive-timeout-sec.
  • ALB retries (global external only): URL map routeAction.retryPolicy (retryConditions, numRetries, perTryTimeout). A custom policy replaces the default one.

Google states that backends behind an Application Load Balancer must use an HTTP keep-alive timeout greater than 600 seconds, and gives 620s as an example (timeouts and retries, passthrough NLB).

Azure​

Azure Load Balancer, Standard (AKS LoadBalancer Service)Application Gateway v2 / AGICApplication Gateway for ContainersFront Door
LayerL4 passthroughL7 proxyL7 proxyL7 proxy (CDN)
Idle timeout4 min (4–100 min)Frontend TCP: 4 min (4–30). Client HTTP/1.1 keep-alive: 120s, fixed. HTTP/2: 180s. Backend: not documented.HTTP idle: 5 min, fixedClient keep-alive: 90s, fixed
When the timeout expiresAKS enables TCP reset by default, so both sides receive an RSTConnection closedConnection closedConnection closed
Request/response timeoutNone20s backend request timeout, which also applies to WebSockets60s request timeout, enforced even while data is streaming30s origin response timeout (16–240s)
Draining / probe failureNo draining. Established flows keep going to a backend whose probe fails until they close or go idle.Connection draining off by default (1–3600s when enabled)
Health probe (AKS)5s interval, 2 failuresAGIC: 30s interval, 3 failures

To configure these values:

  • Azure LB idle timeout: service.beta.kubernetes.io/azure-load-balancer-tcp-idle-timeout: "30" (in minutes).
  • AGIC request timeout: appgw.ingress.kubernetes.io/request-timeout: "60" (seconds).
  • Application Gateway for Containers: HTTPRoute timeouts.request or RoutePolicy routeTimeout.

Microsoft recommends TCP or application-level keepalives with an interval shorter than the idle timeout (TCP reset and idle timeout, AKS load balancer, Application Gateway settings, App Gateway for Containers).

How to align the timeouts​

The recommended values and configuration depend on whether the load balancer is L4 passthrough (AWS NLB, Azure Load Balancer, GCP passthrough NLB) or an L7 proxy (AWS ALB, GCP Application Load Balancer, Azure Application Gateway).

Case 1: Behind an L7 load balancer, Envoy's idle timeout must be longer than the load balancer's​

An L7 load balancer pools connections to Envoy and reuses them. If Envoy closes an idle connection at the moment the load balancer sends a new request on it, that request fails with a 502. The load balancer must always be the side that closes idle connections.

Envoy ClientTrafficPolicy timeout.http.idleTimeout should be greater than the load balancer's backend idle or keep-alive timeout, with some margin.

Load balancerLB backend idle timeoutEnvoy idleTimeout
AWS ALB / Classic ELB60s (or your configured value)Default 1h works. If you set it explicitly, use at least ALB value + 30s.
GCP Application Load Balancer600s, fixedAt least 620s. The default 1h works.
Azure App Gateway for Containers5 minAt least 330s. The default 1h works.
Azure Application GatewayNot documentedKeep the default 1h.

The Envoy default of 1 hour already meets this case. Problems usually start when someone lowers idleTimeout, for example to 300s, which is below the GCP backend keep-alive of 600s.

Case 2: Behind an L4 load balancer, keep connections active or close them before the load balancer forgets them​

L4 load balancers don't pool connections, but they track every flow and drop it after a period of inactivity. The client and Envoy aren't told. The next packet on that connection is reset (AWS NLB, Azure LB) or, after a backend change, sent to the wrong node (GCP). Envoy's default idle timeouts of 1 hour are much longer than the NLB's 350s or Azure's 4 minutes, so idle keep-alive connections, WebSockets and gRPC streams are dropped silently.

Use one or both of these approaches:

  1. Make Envoy close idle connections first. Set the client idleTimeout, and tcp.idleTimeout for TCP and TLS routes, below the load balancer's idle timeout. Envoy then closes cleanly with a FIN or GOAWAY, and clients reconnect without errors.
  2. Keep long-lived connections alive. Enable tcpKeepalive with an idleTime well below the load balancer's idle timeout. TCP keepalive resets the idle timer on the AWS NLB and on Azure LB. HTTP/2 PING keepalive (http2.connectionKeepalive) also works, because L4 load balancers see a PING as traffic.

You can also raise the load balancer's idle timeout, for example to 30 minutes on Azure or up to 6000s on an AWS NLB, and then set Envoy's values relative to the new number.

Load balancerLB idle timeoutEnvoy client idleTimeoutEnvoy tcpKeepalive.idleTime
AWS NLB (TCP)350s300s60s–120s
AWS NLB (TLS listener)350s, fixed300s60s–120s
Azure Load Balancer4 min (240s)200s60s–120s
GCP passthrough NLB60s (re-hashed, not reset)Default is fineOptional. Useful to survive node or backend changes.

Case 3: Timeouts should get shorter the closer they are to the application​

For request timeouts, each hop should allow at least as much time as the hop behind it. The innermost timeout then fires first and returns a clear error from Envoy, with the right response flags and access log entry. Otherwise the load balancer gives up while Envoy is still waiting.

client timeout  ≥  cloud LB request timeout  ≥  Envoy route timeout  ≥  application timeout

Envoy's default route timeout is 15 seconds, which is lower than every cloud default. If any endpoint can take longer than 15 seconds, set the route timeout explicitly and make sure the load balancer allows more:

Load balancerLB request timeoutAction
AWS ALBIdle timeout, 60sRaise idle_timeout.timeout_seconds above the longest Envoy route timeout
GCP Application Load Balancer30s backend service timeoutRaise timeoutSec above the longest Envoy route timeout. For WebSockets, use a large value such as 3600s.
Azure Application Gateway20sRaise request-timeout above the longest Envoy route timeout. This also applies to WebSockets.
Azure App Gateway for Containers60s, also applied to streamingRaise it or set it to 0s for SSE or long downloads
Azure Front Door30sUp to 240s
L4 load balancersNoneOnly Case 2 applies

For streaming, SSE and WebSocket routes, also check the stream idle timeout (5 minutes by default). Raise it with BackendTrafficPolicy timeout.http.streamIdleTimeout, or set a long route timeout, which raises the route's stream idle timeout with it.

Case 4: Envoy's upstream idle timeout must be shorter than the backend's keep-alive​

Case 1 applies again on the next hop, with Envoy as the client. Envoy pools connections to each backend and keeps an idle one for 1 hour by default. Most application servers close idle keep-alive connections much sooner: Node.js after 5s, Go net/http servers when IdleTimeout is set, gunicorn after 2s, and many frameworks after 30–120s. When Envoy sends a request on a connection the backend is closing, the request fails with a 503 and response flag UC (upstream connection termination). These errors are intermittent and scale with traffic, so they're easy to miss.

Set BackendTrafficPolicy timeout.http.connectionIdleTimeout below the shortest keep-alive timeout of the backends it applies to. One policy for the whole Gateway only works if every backend has a longer keep-alive than the value you choose. For backends with very short keep-alives, either:

  • attach a separate BackendTrafficPolicy to their routes with a lower value, or
  • raise the keep-alive timeout in the application. This is usually better, because it avoids reconnecting on almost every request.

A retry on reset-before-request (see below) hides most of the remaining UC errors, but it doesn't prevent them. Fix the timeouts first.

Retries across layers​

Retries and timeouts interact. Every retry adds another timeout's worth of latency, and when several layers retry, the attempts multiply. A cloud load balancer that retries twice in front of an Envoy that retries twice can send the same request to a backend up to nine times.

Retry in one layer, preferably Envoy: it knows whether any bytes reached the backend, so it can retry safely, and its access logs show every attempt. Limit the cloud load balancer to retrying connection-level failures, or nothing.

Envoy​

Envoy Gateway doesn't configure any retries by default. For a retry that's safe for any method, including POST, use triggers that only fire when the request didn't reach the backend:

TriggerFires whenSafe for non-idempotent requests?
connect-failureThe upstream connection couldn't be establishedYes
reset-before-requestThe connection was reset before the request was sent, for example on a pooled connection the backend just closedYes
refused-streamThe backend refused the HTTP/2 stream (REFUSED_STREAM)Yes
resetThe connection was reset at any point, including after the backend received the requestNo
5xx, gateway-error, retriable-status-codesThe backend or Envoy returned an error statusNo. The backend may have processed the request.
warning

Adding a retry block without retryOn applies Envoy Gateway's default triggers. They include retriable-status-codes for 503, which replays requests the backend may already have processed. Always list the triggers explicitly.

Retries count against the route timeout, which covers all attempts. Use perRetry.timeout to bound each attempt.

Google Cloud Application Load Balancer​

The global external Application Load Balancer retries by default: when no retry policy is configured, it retries body-less requests (such as GET) once on any 502, 503 or 504. That has two side effects:

  • A request that times out waits for the backend service timeout twice. With the 30s default, a GET that hits a 504 costs the client 60s.
  • A GET that reached the backend and failed afterwards runs twice. If it has side effects, such as firing webhooks, they happen twice.

To change this, set routeAction.retryPolicy on the URL map. A custom policy replaces the default one. Restrict retryConditions to failures where the request didn't reach Envoy, such as connect-failure and refused-stream, and let Envoy handle the rest.

info

In Google Cloud, numRetries is the total number of attempts, including the first one. numRetries: 2 means one retry, and numRetries: 1 disables retries. In Envoy, numRetries counts retries only, excluding the first attempt.

When no custom retry policy exists, the load balancer uses the backend service timeout for each attempt and ignores routeAction.timeout. The classic Application Load Balancer always retries GET requests that fail before response headers, as long as at least 80% of backends are healthy, and this can't be changed (timeouts and retries).

Draining during rollouts and scale-down​

When an Envoy pod terminates, the load balancer needs time to notice before Envoy stops accepting connections:

  • minDrainDuration should be at least as long as the load balancer needs to mark the target unhealthy (health check interval × unhealthy threshold), plus some margin. The 10s default is too short for most cloud load balancers. Examples: AWS NLB with default target group settings is 30s × 2 = 60s; GKE NEG is 15s × 2 = 30s; Azure LB is 5s × 2 = 10s.
  • drainTimeout should cover your longest normal requests. terminationGracePeriodSeconds is adjusted to drainTimeout + 300s automatically.
  • AWS NLB: the default deregistration delay of 300s doesn't close connections when it expires. Set deregistration_delay.timeout_seconds to roughly match drainTimeout, and consider enabling deregistration_delay.connection_termination.enabled.
  • Azure Load Balancer never drains or resets established flows when a probe fails. Envoy closing connections during drain (with Connection: close or GOAWAY) is what moves clients to healthy pods. Make sure drainTimeout is long enough for that.

Example configurations​

The examples below apply to a Gateway named eg. Adjust names and namespaces to match your environment.

AWS NLB or Azure Load Balancer (L4)​

apiVersion: gateway.envoyproxy.io/v1alpha1
kind: ClientTrafficPolicy
metadata:
name: cloud-lb-timeouts
namespace: default
spec:
targetRefs:
- group: gateway.networking.k8s.io
kind: Gateway
name: eg
timeout:
http:
idleTimeout: 300s # below NLB 350s; use 200s for Azure LB (240s)
tcp:
idleTimeout: 300s # for TCPRoute/TLSRoute listeners
tcpKeepalive:
idleTime: 60s # keep long-lived connections alive at L4
interval: 30s
probes: 3

AWS ALB or GCP Application Load Balancer (L7)​

apiVersion: gateway.envoyproxy.io/v1alpha1
kind: ClientTrafficPolicy
metadata:
name: cloud-lb-timeouts
namespace: default
spec:
targetRefs:
- group: gateway.networking.k8s.io
kind: Gateway
name: eg
timeout:
http:
idleTimeout: 3600s # must stay above the LB backend keep-alive (ALB 60s, GCP 600s)

Request timeouts​

Set the route timeout to what the application needs, and raise the load balancer's timeout above it:

apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
name: backend
spec:
parentRefs:
- name: eg
rules:
- backendRefs:
- name: backend
port: 3000
timeouts:
request: 50s # Envoy default is 15s; keep below the LB timeout (e.g. ALB 60s)

Upstream idle timeout and safe retries​

Apply a Gateway-wide default, then override it for backends with very short keep-alives:

apiVersion: gateway.envoyproxy.io/v1alpha1
kind: BackendTrafficPolicy
metadata:
name: upstream-defaults
namespace: default
spec:
targetRefs:
- group: gateway.networking.k8s.io
kind: Gateway
name: eg
timeout:
http:
connectionIdleTimeout: 25s # below the shortest backend keep-alive covered by this policy
retry:
numRetries: 1 # Envoy counts retries, excluding the first attempt
retryOn:
triggers: # only fire when the request didn't reach the backend
- connect-failure
- reset-before-request
- refused-stream
---
apiVersion: gateway.envoyproxy.io/v1alpha1
kind: BackendTrafficPolicy
metadata:
name: short-keepalive-backend
namespace: default
spec:
targetRefs:
- group: gateway.networking.k8s.io
kind: HTTPRoute
name: node-app # backend closes idle connections after 5s
mergeType: StrategicMerge # keep the Gateway-level retry policy
timeout:
http:
connectionIdleTimeout: 4s

Without mergeType, a route-level BackendTrafficPolicy replaces the Gateway-level one for that route, and the route loses the retry policy.

Google Cloud ALB retry policy​

Limit the global external Application Load Balancer to connection-level retries in the URL map, and leave the rest to Envoy:

defaultRouteAction:
retryPolicy:
retryConditions:
- connect-failure
- refused-stream
numRetries: 2 # total attempts in Google Cloud: one retry

Load balancer settings and graceful shutdown​

Set load balancer annotations and shutdown timing on the EnvoyProxy resource referenced by your GatewayClass. When you install TEG with Helm, set them under config.envoyProxy in the teg-envoy-gateway-helm values. The example below is for an AWS NLB:

apiVersion: gateway.envoyproxy.io/v1alpha1
kind: EnvoyProxy
metadata:
name: teg-envoy-proxy
namespace: envoy-gateway-system
spec:
shutdown:
minDrainDuration: 60s # >= LB health check interval x unhealthy threshold
drainTimeout: 90s
provider:
type: Kubernetes
kubernetes:
envoyService:
annotations:
service.beta.kubernetes.io/aws-load-balancer-type: external
service.beta.kubernetes.io/aws-load-balancer-nlb-target-type: ip
service.beta.kubernetes.io/aws-load-balancer-target-group-attributes: >-
deregistration_delay.timeout_seconds=90,deregistration_delay.connection_termination.enabled=true
# For Azure instead:
# service.beta.kubernetes.io/azure-load-balancer-tcp-idle-timeout: "30"

Troubleshooting with access logs​

Envoy access logs show which hop closed the connection:

SymptomLikely cause
502 at the cloud load balancer, with no matching Envoy access log entryEnvoy closed an idle connection that the load balancer reused. Raise Envoy's idleTimeout (Case 1).
504 from Envoy with response flag UTThe route timeout expired. The default is 15s (Case 3).
504 at the load balancer while the Envoy log shows a long %DURATION% or no entry yetThe load balancer's request timeout is shorter than Envoy's route timeout (Case 3).
Clients see connection reset on reused or long-lived connectionsAn L4 load balancer idle timeout expired (Case 2).
Response flag SIThe stream idle timeout expired (5 min by default).
503 with response flag UCThe backend application closed a pooled connection that Envoy reused. Set BackendTrafficPolicy timeout.http.connectionIdleTimeout below the application's keep-alive timeout (Case 4).
GET requests reach the backend twice, or side effects (such as webhooks) fire twiceThe Google Cloud global external ALB default retry replayed a request that got a 502, 503 or 504. Set a custom retryPolicy (Retries across layers).
A 504 takes twice the configured timeout (for example 60s with a 30s timeout)A retry layer waited for the full timeout on each attempt, such as the Google Cloud ALB default retry. Restrict retries to connection failures, or set perRetry.timeout in Envoy.
Several Envoy access log entries for one client request, or %UPSTREAM_REQUEST_ATTEMPT_COUNT% above 1Retries in more than one layer multiplied. Keep retries in Envoy and limit the load balancer to connection-level failures.