n8n Error Workflows That Actually Alert: Deduplication, Stale Locks and the Friday Email
Part 6 of 6 in the series The Follow-Up Machine
- Overview, Automated CRM Follow-Up That Tells You When It Breaks
- Architecture, An n8n Workflow Example Built to Survive Retries: The Follow-Up Machine Architecture
- Register, EspoCRM Webhooks in n8n: The Event Name, the Signature and the Stage That Did Not Exist
- Evaluate, An Hourly n8n Workflow That Does Not Overlap Itself: Locks, Merge Barriers and Quiet Runs
- Enrich, AI Summaries of CRM Emails with n8n and Mistral: The Prompt, the Parser and the Comma
- Alerts, n8n Error Workflows That Actually Alert: Deduplication, Stale Locks and the Friday Email
All 6 parts
- Overview, Automated CRM Follow-Up That Tells You When It Breaks
- Architecture, An n8n Workflow Example Built to Survive Retries: The Follow-Up Machine Architecture
- Register, EspoCRM Webhooks in n8n: The Event Name, the Signature and the Stage That Did Not Exist
- Evaluate, An Hourly n8n Workflow That Does Not Overlap Itself: Locks, Merge Barriers and Quiet Runs
- Enrich, AI Summaries of CRM Emails with n8n and Mistral: The Prompt, the Parser and the Comma
- Alerts, n8n Error Workflows That Actually Alert: Deduplication, Stale Locks and the Friday Email
This is the last part of the follow-up machine series. Parts 3 to 5 covered the workflows that do the work. This part covers the two that tell me whether the work is happening: “Followup 0: Alerts” and “Followup 3: Report”.
The rule behind both is the one from the start of the project: a system that silently stops is worse than no system, because you keep trusting it.
An inactive error workflow did not fire on n8n 2.31.7
n8n lets a workflow name an error workflow in its settings. When an execution fails, n8n starts that error workflow with the failure details. All four working workflows point to “Followup 0: Alerts”.
The first real failure was a scheduled run of Evaluate on 4 August. The error workflow was set correctly. It produced no execution at all.
n8n error workflows. n8n’s documentation says a workflow that starts with an Error Trigger does not need to be activated. On the version I run, 2.31.7, an inactive one did not fire.
Activating it fixed it. I have not checked newer versions. There is no downside: an Error Trigger has no schedule and no webhook, so an active error workflow does nothing until something fails.
That is the kind of bug this whole part is about.
The alarm was installed, configured and silent.
Alert once, not every hour
The alert workflow is short:
On Error (Error Trigger)
→ Log + dedup (Postgres)
├─→ Should alert? ──yes──→ Dry run? ──no──→ Send alert (Outlook)
│ └─no──→ No alert (deduped)
└─→ Release stale claim (Postgres)
Deduplication lives in the same query that logs the failure:
insert into run_log (workflow, run_id, started_at, finished_at, status, error_count, notes)
values (
$1::jsonb ->> 'workflow', gen_random_uuid(), now(), now(), 'failed', 1,
jsonb_build_object(
'error', $1::jsonb ->> 'error',
'alerted', not exists (
select 1 from run_log
where workflow = $1::jsonb ->> 'workflow'
and status = 'failed'
and (notes ->> 'alerted')::boolean
and started_at > now() - interval '6 hours'
)
)
)
returning (notes ->> 'alerted')::boolean as should_alert;
Every failure is recorded. An email goes out only if no alert went out for the same workflow in the last six hours.
The reason is practical. Evaluate runs 13 times a day, and Enrich another 13. An expired CRM credential fails every one of those runs. Without deduplication that is up to 26 alert emails a day, and the reasonable human response is a mail filter. After that, the next real failure lands in the filter too.
The parameter is one JSON object, for the reason from part 5: error messages contain commas, and the n8n Postgres node splits its parameter string on them.
Releasing the lock of a crashed run
Part 2 described the run claim: a run_log row with status running that blocks other runs for 55 minutes. A run that finishes normally releases it. A run that crashes never reaches its last node, so its claim stays.
The consequence: any run of the same workflow started within the next 55 minutes found the claim and stopped without processing anything. By the next scheduled run, 60 minutes later, the claim had just expired. A manual rerun straight after fixing the problem had not: the failed run still held its claim, so the immediate retry was skipped and only the run after that got through. During acceptance testing, with runs started by hand minutes apart, one failure regularly cost an extra attempt.
The error workflow knows which workflow died, so it now releases the claim:
update run_log
set status = 'failed',
finished_at = now(),
notes = coalesce(notes, '{}'::jsonb) || jsonb_build_object('released_by', 'alerts')
where workflow = $1::jsonb ->> 'workflow'
and coalesce($1::jsonb ->> 'workflow', '') <> ''
and status = 'running'
returning id;
It is wired parallel to the alert branch, so it runs even when the alert email is suppressed by deduplication. It is safe because n8n ends a failed execution before the error workflow starts: nothing live gets released. The 55-minute window stays as the backstop for the one case the error workflow cannot see, n8n itself dying mid-run.
The report that always sends
“Followup 3: Report” has two schedules on one trigger:
| Schedule | Content | Sends without items? |
|---|---|---|
| Tuesday to Friday, 08:45 | Follow-ups due today, with context and suggested next step | No |
| Friday, 15:45 | Deals that moved, deals that stalled, next week, closed this week, plus a health line | Yes |
The Friday email is sent even when nothing happened. An empty digest is not noise. It is proof that the machine is alive, delivered without me having to look for it. That holds as long as the run itself gets through: as the section further down shows, a database outage means it sends nothing at all.
Its footer is the health check:
export function healthFooter(h: Health): { line: string; critical: boolean } {
if (!h.lastOkWf2At || h.now.getTime() - h.lastOkWf2At.getTime() > 24 * 3_600_000) {
return {
critical: true,
line: `⚠ WF2 has not run successfully in the last 24 hours (last ok: ${h.lastOkWf2At?.toISOString() ?? 'never'}). ${h.openWatches} open watches.`,
};
}
return {
critical: false,
line: `Health: last WF2 ok ${h.lastOkWf2At.toISOString()} · ${h.partialRunsThisWeek} partial run(s) this week · ${h.openWatches} open watches`,
};
}
When the critical flag is set, that line moves from the bottom of the email to the top.
The report also refuses to send a partial digest. If a query fails, the run fails and the error workflow takes over. A digest with a section silently missing teaches you to distrust the email, and then it stops doing its job.
The quiet morning that looked like a dead workflow
The daily branch had its own version of the empty-run bug from part 4. On a morning with nothing due, the compose step returned zero items. The send nodes were skipped, which was correct. But the Merge node in front of “Log run” was waiting for the send path, and it starved. No run_log row was written.
So a healthy, quiet morning and a dead workflow looked identical in run_log. That is precisely what the health checks are supposed to tell apart.
The fix is the same pattern as in part 4. Compose returns one { skip: true } item instead of nothing, and an IF node, “Anything to send?”, routes it past the send nodes straight to the Merge node. “Log run” then counts sent items with a guard, because referencing a node that did not execute throws. The guard looks like this:
{{ $('Send').isExecuted ? $items('Send').length : 0 }}
What the alert path still cannot report
Look at the alert workflow again. The first thing it does is write to Postgres. Both of its Postgres nodes use the same database, and the same credential, as every other workflow in the machine. Neither has an error output. The alert workflow itself has no error workflow.
So if Postgres is down, or its credential is broken, the chain goes like this: Evaluate fails at its first database node, the error workflow starts, “Log + dedup” fails after three tries, meaning one call and two retries, and the error workflow stops. “Send alert” never runs. The Friday report also needs Postgres, and by design sends nothing rather than a partial digest.
The gap. The one failure the architecture calls fatal is the one the alert path cannot announce. The only signal left is an absence: no Friday email.
I found this by reading the workflow while writing this post, not through an outage. The fix would be small: “Log + dedup” gets an error output that goes straight to “Send alert”, with a subject that says deduplication is unavailable. That gives up the one-alert-every-six-hours limit during a database outage, because the limit lives in exactly the query that fails then. In exchange, the alert path gets a chance to send anything at all. For the failure that stops everything I think that is the right trade. It is not built as of this post.
The general lesson is the one this series keeps arriving at.
An alarm that depends on the thing it watches may not be able to report that thing failing.
Smaller things that cost real time
- Workflow names are load-bearing. Alert subjects and
run_loglabels are derived from the display name, “Followup 2: Evaluate” and so on. Renaming a workflow silently breaks the claim release, which parses the number out of that name. One note on spelling: in the running system a long dash separates the number from the word, and this series writes a colon throughout for readability. - Node IDs must be real UUIDs. The first workflows were written with short hand-made node IDs like
w4. The n8n editor finds a node from a link by ID prefix, so openingw4landed onw401. Every node now has acrypto.randomUUID()ID. - Document inside n8n, not only beside it. All five workflows had no description and no sticky notes, so the editor showed five names and canvases of up to 40-odd nodes with no statement of intent. Each now has a description and a “Purpose” sticky note, generated from the same source as the written documentation.
- Generating documentation finds stale configuration. The generator that writes those descriptions surfaced two leftover one-off test schedules in the saved workflow files, which had already been reverted in the live instance.
What to take from this series
- Decide which system holds the authoritative data before you build anything.
- Make every workflow safe to run twice.
- Test the run where nothing happens.
- Test with data that looks like real data.
- Make the machine prove it is alive on a schedule, even when it has nothing to say.
- Check that the alarm does not depend on the thing it watches.
The series started with the overview of what the machine does. If a process in your business goes quiet the same way, tell me about it.