n8n error handling: retries, alerts and recovery.
Plan recoverable failures instead of letting automation silently stop.
Classify failures
Separate temporary errors such as rate limits from permanent errors such as invalid input or revoked permissions.
Use bounded retries
Set a small retry limit, use delays for transient failures and alert a named owner after the limit. Never retry a financial side effect blindly.
Create an exception path
Record the run ID, failed step, sanitized error and source record. Avoid putting credentials or sensitive payloads in chat alerts.
Verify recovery
After retry or manual repair, verify the target state and mark the run recovered only when the expected result is present.
Exercise the failure path
Test provider timeout, authorization failure, malformed response and exhausted retries in a test workspace.
Use the toolkit
Turn this advice into a structured workflow, then test it with synthetic data before production.