Reliability ยท Practical guide

n8n error handling: retries, alerts and recovery.

Plan recoverable failures instead of letting automation silently stop.

Principle: document the expected outcome, the failure path and the proof of completion before connecting live systems.

Classify failures

Separate temporary errors such as rate limits from permanent errors such as invalid input or revoked permissions.

Use bounded retries

Set a small retry limit, use delays for transient failures and alert a named owner after the limit. Never retry a financial side effect blindly.

Create an exception path

Record the run ID, failed step, sanitized error and source record. Avoid putting credentials or sensitive payloads in chat alerts.

Verify recovery

After retry or manual repair, verify the target state and mark the run recovered only when the expected result is present.

Exercise the failure path

Test provider timeout, authorization failure, malformed response and exhausted retries in a test workspace.

Use the toolkit

Turn this advice into a structured workflow, then test it with synthetic data before production.

Build a workflow โ†— Browse agency templates