connect ETIMEDOUTon outbound connections from Worker service
15 days ago
Our n8n Worker service intermittently fails to open TCP connections to a single
external destination (AWS RDS PostgreSQL, port 5432). The connection attempt
times out during the TCP handshake — connect ETIMEDOUT — with no
ECONNREFUSED and no database-level error. Packets appear to be silently
dropped.
Failures arrive in bursts, several within the same minute, interleaved with
successful connections to the same destination. Roughly 2% of our workflow
executions are affected.
We have investigated on our side and ruled out the destination firewall, the
database connection limit, our application connection pool, DNS resolution and
service resources. The behaviour correlates only with outbound connection
volume: the service that opens many short-lived connections fails, while
services on the same egress path that open few connections never do.
We would like your help identifying what is happening on the outbound network
path, and whether dedicated outbound IPs would help.
Environment
- Project:
n8n production(fd4b3b05-8dda-4bdf-a98d-1e72f588d86f) - Environment:
production(87ed90e9-fb26-4ad4-8a2b-ae17a88506aa) - Region:
europe-west4-drams3a
Affected service:
-
Worker (
3af14a9a-b457-4dfa-becc-3ee58f5f08d8), imagen8nio/n8n,7 replicas, deployment
021f3162-4fe9-4149-80be-8ca8d404a529(redeployed today at 10:29 UTC)
Unaffected services, same project and same egress path:
- Primary (
62c87730-24b3-4f38-83ad-a759002b0cf7), 1 replica - Webhook processor (
bbd118b0-d451-4509-a672-9dd4763fb1fd), 2 replicas
Outbound networking
Static Outbound IPs are enabled, type: Shared, High Availability:
208.77.244.242152.55.184.240152.55.184.241
Destination
-
54.155.216.50:5432— AWS Aurora PostgreSQL cluster endpoint, publiclyaccessible,
eu-west-1 -
This is the only affected destination. Every failure targets this single
IP and port.
Sample logs
From the Worker deploy logs (UTC):
12:29:13 connect ETIMEDOUT 54.155.216.50:5432
12:29:20 connect ETIMEDOUT 54.155.216.50:5432
12:30:02 connect ETIMEDOUT 54.155.216.50:5432
12:30:02 connect ETIMEDOUT 54.155.216.50:5432
12:30:50 connect ETIMEDOUT 54.155.216.50:5432
12:31:25 connect ETIMEDOUT 54.155.216.50:5432
12:31:28 connect ETIMEDOUT 54.155.216.50:5432
12:31:45 connect ETIMEDOUT 54.155.216.50:5432
12:32:08 connect ETIMEDOUT 54.155.216.50:5432
12:32:14 connect ETIMEDOUT 54.155.216.50:5432
12:32:45 connect ETIMEDOUT 54.155.216.50:5432
12:32:48 connect ETIMEDOUT 54.155.216.50:5432
12:32:48 connect ETIMEDOUT 54.155.216.50:5432Earlier the same day, 8 failures within a 40-second window
(12:02:27 → 12:03:07).
The service or the route. Primary and Webhook processor use the identical
egress path and the same static IPs, and have logged zero ETIMEDOUT since
10:35 UTC. The difference is volume: Worker executes all workflow steps and
opens far more short-lived connections than the other two.
Questions
-
Are there any limits on outbound connections from Shared static IPs —
concurrency, rate, or per-destination — that could cause connection attempts
to be dropped rather than refused?
-
All of our outbound traffic targets one destination IP and port. Is that
pattern known to cause problems on shared egress?
Impact
These are production automations handling mission assignment, SMS dispatch and
data synchronisation. Each failed execution is a business event that does not
happen. We have added retries on the affected steps as a stopgap, but we would
like to resolve the underlying cause.
1 Replies
Status changed to Awaiting Railway Response Railway • 15 days ago
15 days ago
The flow records from your Worker service show zero dropped packets on egress, and every other external destination (Airtable, Supabase, your own ELB endpoints) is completing TCP handshakes and exchanging data normally. The connections to 54.155.216.50:5432 that fail are sending a SYN that leaves our network and is never answered - the destination is silently discarding those specific connection attempts.
Because other outbound connections from the same containers succeed at the same time, this rules out a host-level egress fault. The packets left, and the Aurora endpoint chose not to respond to some of them.
The most likely cause is a connection-rate or per-source-IP limit on the Aurora side. Your three static IPs are shared with other workloads, so the cumulative connection rate to that specific destination from those addresses may be higher than what your Worker alone sends. AWS security groups and Aurora connection throttling both operate silently (drop, not reject) at high connection rates, which produces exactly the ETIMEDOUT-with-no-ECONNREFUSED pattern you described.
To your questions: there are no per-destination connection limits imposed on our egress side. Dedicated (unshared) outbound IPs are not available on any plan. The resolution path is on the destination side - allowlisting your three static IPs in the Aurora security group if they are not already there, and checking whether AWS is applying any network-level rate limiting to inbound connections on that instance.
Status changed to Awaiting User Response Railway • 15 days ago
Status changed to Solved sam-a • 15 days ago