Intermittent
connect ETIMEDOUT
on outbound connections from Worker service
otoqi
PROOP

15 days ago

Our n8n Worker service intermittently fails to open TCP connections to a single

external destination (AWS RDS PostgreSQL, port 5432). The connection attempt

times out during the TCP handshake — connect ETIMEDOUT — with no

ECONNREFUSED and no database-level error. Packets appear to be silently

dropped.

Failures arrive in bursts, several within the same minute, interleaved with

successful connections to the same destination. Roughly 2% of our workflow

executions are affected.

We have investigated on our side and ruled out the destination firewall, the

database connection limit, our application connection pool, DNS resolution and

service resources. The behaviour correlates only with outbound connection

volume: the service that opens many short-lived connections fails, while

services on the same egress path that open few connections never do.

We would like your help identifying what is happening on the outbound network

path, and whether dedicated outbound IPs would help.

Environment

  • Project: n8n production (fd4b3b05-8dda-4bdf-a98d-1e72f588d86f)
  • Environment: production (87ed90e9-fb26-4ad4-8a2b-ae17a88506aa)
  • Region: europe-west4-drams3a

Affected service:

  • Worker (3af14a9a-b457-4dfa-becc-3ee58f5f08d8), image n8nio/n8n,

    7 replicas, deployment 021f3162-4fe9-4149-80be-8ca8d404a529

    (redeployed today at 10:29 UTC)

Unaffected services, same project and same egress path:

  • Primary (62c87730-24b3-4f38-83ad-a759002b0cf7), 1 replica
  • Webhook processor (bbd118b0-d451-4509-a672-9dd4763fb1fd), 2 replicas

Outbound networking

Static Outbound IPs are enabled, type: Shared, High Availability:

  • 208.77.244.242
  • 152.55.184.240
  • 152.55.184.241

Destination

  • 54.155.216.50:5432 — AWS Aurora PostgreSQL cluster endpoint, publicly

    accessible, eu-west-1

  • This is the only affected destination. Every failure targets this single

    IP and port.

Sample logs

From the Worker deploy logs (UTC):

12:29:13  connect ETIMEDOUT 54.155.216.50:5432
12:29:20  connect ETIMEDOUT 54.155.216.50:5432
12:30:02  connect ETIMEDOUT 54.155.216.50:5432
12:30:02  connect ETIMEDOUT 54.155.216.50:5432
12:30:50  connect ETIMEDOUT 54.155.216.50:5432
12:31:25  connect ETIMEDOUT 54.155.216.50:5432
12:31:28  connect ETIMEDOUT 54.155.216.50:5432
12:31:45  connect ETIMEDOUT 54.155.216.50:5432
12:32:08  connect ETIMEDOUT 54.155.216.50:5432
12:32:14  connect ETIMEDOUT 54.155.216.50:5432
12:32:45  connect ETIMEDOUT 54.155.216.50:5432
12:32:48  connect ETIMEDOUT 54.155.216.50:5432
12:32:48  connect ETIMEDOUT 54.155.216.50:5432

Earlier the same day, 8 failures within a 40-second window

(12:02:27 → 12:03:07).

The service or the route. Primary and Webhook processor use the identical

egress path and the same static IPs, and have logged zero ETIMEDOUT since

10:35 UTC. The difference is volume: Worker executes all workflow steps and

opens far more short-lived connections than the other two.

Questions

  1. Are there any limits on outbound connections from Shared static IPs —

    concurrency, rate, or per-destination — that could cause connection attempts

    to be dropped rather than refused?

  2. All of our outbound traffic targets one destination IP and port. Is that

    pattern known to cause problems on shared egress?

Impact

These are production automations handling mission assignment, SMS dispatch and

data synchronisation. Each failed execution is a business event that does not

happen. We have added retries on the affected steps as a stopgap, but we would

like to resolve the underlying cause.

Solved

1 Replies

Status changed to Awaiting Railway Response Railway • 15 days ago


sam-a
EMPLOYEE

15 days ago

The flow records from your Worker service show zero dropped packets on egress, and every other external destination (Airtable, Supabase, your own ELB endpoints) is completing TCP handshakes and exchanging data normally. The connections to 54.155.216.50:5432 that fail are sending a SYN that leaves our network and is never answered - the destination is silently discarding those specific connection attempts.

Because other outbound connections from the same containers succeed at the same time, this rules out a host-level egress fault. The packets left, and the Aurora endpoint chose not to respond to some of them.

The most likely cause is a connection-rate or per-source-IP limit on the Aurora side. Your three static IPs are shared with other workloads, so the cumulative connection rate to that specific destination from those addresses may be higher than what your Worker alone sends. AWS security groups and Aurora connection throttling both operate silently (drop, not reject) at high connection rates, which produces exactly the ETIMEDOUT-with-no-ECONNREFUSED pattern you described.

To your questions: there are no per-destination connection limits imposed on our egress side. Dedicated (unshared) outbound IPs are not available on any plan. The resolution path is on the destination side - allowlisting your three static IPs in the Aurora security group if they are not already there, and checking whether AWS is applying any network-level rate limiting to inbound connections on that instance.


Status changed to Awaiting User Response Railway • 15 days ago


Status changed to Solved sam-a • 15 days ago


Welcome!

Sign in to your Railway account to join the conversation.

Loading...