Why your n8n AI agent's HTTP calls time out (and how to fix them)
AI agent workflows make far more HTTP calls than classic automations: model APIs, tool endpoints, webhooks, vector stores. When those calls are slow, n8n's defaults turn intermittent slowness into failed executions.
What the default behavior is
Every HTTP Request node in n8n has an Options → Timeout setting, in milliseconds. If you never set it, the node uses its built-in default of 10,000 ms (10 seconds). A model API that takes 12 seconds to stream a response will fail the node with an ETIMEDOUT-style error — even though the API was working fine.
There is also a workflow-level execution timeout (workflow Settings → Timeout Workflow). A node timeout that is too generous can push the whole run into the workflow timeout instead. Set both deliberately.
Fix 1: set an explicit timeout per node
Open the HTTP Request node → Options → add Timeout. Use 10,000 ms for fast health checks, 30,000–60,000 ms for slow model or embedding APIs. In exported workflow JSON this appears as:
{
"name": "Call LLM API",
"type": "n8n-nodes-base.httpRequest",
"parameters": {
"method": "POST",
"url": "https://api.example.com/v1/chat",
"options": { "timeout": 60000 }
},
"onError": "continueErrorOutput"
}
Fix 2: add retries — but only for idempotent calls
Select the node → Settings tab → enable Retry On Fail, then set Max Tries and Wait Between Tries. In JSON:
"retryOnFail": true,
"maxTries": 3,
"waitBetweenTries": 2000
Only retry requests that are safe to repeat (GETs, idempotent POSTs). Retrying a payment or record-creation call can double-apply the side effect.
Fix 3: decide what a failure means
Every node has an onError behavior. The default, stopWorkflow, halts the entire execution on the first failed call. For agent workflows, continueErrorOutput routes the failure to the node's error output so a downstream node can fall back, log, or escalate — the agent keeps working instead of dying.
Diagnose it in the exported JSON
Export the workflow (⋮ menu → Download) and run this jq one-liner to find every HTTP node missing a timeout, error handling, or retry:
jq -r '.nodes[]
| select(.type == "n8n-nodes-base.httpRequest")
| "\(.name): timeout=\(.parameters.options.timeout // "NOT SET")
onError=\(.onError // "stopWorkflow (default)")
retryOnFail=\(.retryOnFail // false)"' workflow.json
Expected output for a healthy workflow:
Call LLM API: timeout=60000 onError=continueErrorOutput retryOnFail=true
Fetch embeddings: timeout=30000 onError=continueErrorOutput retryOnFail=true
Any line showing NOT SET or stopWorkflow (default) is a node that will fail loudly the first time an API is slow.
Common failures
- No timeout set on slow LLM calls. The 10-second default kills calls that just needed 15 seconds.
- No
onErrorhandling. One flaky webhook stops a 40-node agent run. - Retry storms. Retrying a non-idempotent POST, or retrying aggressively against a rate-limited API, turns one failure into many — and can get the API key throttled.
- Timeout set too high. A 5-minute node timeout on every node means the workflow-level execution timeout fires first, and the error message points at the wrong place.
Go further
The free n8n-workflow-linter checks exported workflow JSON for exactly these smells — missing timeouts, missing error handling, and more. For the full treatment (10 importable reliability workflows, a failure-simulation eval harness, readiness scoring, and a 37-page production playbook), see the n8n Production AI Agent Reliability Kit ($99, one-time).
Built by Payload
Payload builds practical software that makes AI, automation, and business infrastructure safer, cleaner, more reliable, and easier to ship. Support: kylers.partners@gmail.com