Stop duplicate side effects on webhook retries

had a client flow create 3 hubspot notes for one form submit because the sender retried.

fix that stuck for me:

grab a stable id (their idempotency key, or hash of external_id + event type). before any create/update node, check a small store for that key. if it’s there, return 200 and exit. if not, write the key first, then do the side effects.

also label which nodes write vs read in the sticky note. future you will forget.

not perfect for every case but killed the duplicate notes.

@blessoftware - the atomic step is the right fix, and I’d extend it on three edges that bite in production.

Key choice. “Their idempotency key” only works when the sender actually provides a stable one, so check the delivery headers before trusting them: some providers send a stable event id that survives retries, others rotate per attempt (Slack rotates its request timestamp and signature on every retry, so any header-derived key becomes a brand-new key each time). The robust fallback is a hash of the canonical body: an identical payload redelivered is a retry, while a similar-but-new payload (the user genuinely submits the form twice) is a new event. Hashing external_id + event type conflates those two cases and will eventually drop a legitimate second submit.

TTL is a guarantee window, not a housekeeping setting. 86400 really means “after 24 hours I can no longer tell a retry from a first delivery.” Size it off the sender’s documented redelivery schedule rather than tidiness: if the sender retries for up to 3 days, any TTL under 3 days plus margin leaves a hole at the end where a late retry sails through as fresh. Cheap insurance: TTL = maximum redelivery window + a day.

Seen is not done. Both of these snippets record “we started processing this key” but nothing records that the HubSpot create actually succeeded. If the workflow dies after the INCR/INSERT but before the note is written, every later retry gets stopped as a duplicate - and now you don’t have a duplicate, you have silently missing data, which is worse. The store needs a status on the record (claimed → done/failed) and a small sweeper that re-queues anything stuck in claimed past some age.

That last point is exactly what Respond to Immediately trades on: once you 200 early, the sender will never redeliver, so your dedup record becomes the only evidence the event ever arrived. Reasonable trade - but it moves the durability requirement from the sender onto your store, so the store has to earn it.

Curious what you’ve seen on the TTL side in practice: do senders actually exceed their documented redelivery windows, or is documented-window-plus-margin holding up for you?

@InvestigatorSuper216 Top of the day to you!
The solutions provide by @blessoftware and @InvestigatorSuper216 are highlty accurate, technically sound, and align perfectly with distributed system best practices.

Fact-Check and Technical Highlights

Atomic oper operations @blessoftware is completely correct. A standard “check-then-write” pattern creates a race condition (TOCTOU). Using Redis INCR or PostgreSQL INSERT ... ON CONFLICT DO NOTHING guarantees atomicity.

The output Trap: In n8n, disabling Always Output Data on a database node ensures that when a conflict returns 0 rows, the workflow branc natively terminates there.

webhook Response Selection: Setting the n8n webhook node response mode to "Respond Immediately " is the best practice to prevent sender timeouts.

"Seen is not Done : @iwasinnam21 accurately points out the "Dual-write " flaw.
Locking the key before calling HubSpot risks silent data loss if the workflow fails mid-execution.

Here are n8n native alternatives…
Built-in n8n features to simplify the setup:

n8n Data Tables (No External DB needed): instead of hosting a separate Redis/postgres instance, use built in n8n data tables “Data tables | Build | n8n Docs” Set a unique constraint on your custom key column . A duplicate insertion naturally throws an error fails to append, halting the branch.

Remove Duplicates Node: For non-concurrent duplicates(spaced minutes or hours apart), use the native n8n Remove Duplicates node"Remove Duplicates | Nodes | n8n Docs". set the operation to compare against previous executions using a hashed payload string.

Error Routing: To solve the “Seen is not done” problem natively , use n8n error handling . if HubSpot fails , route the error to a node that deletes or clears the locked key from your store so it can be retried later.

One nuance I’d add: “write the key first” prevents the duplicate side effect, but a simple seen/not-seen flag can create the opposite failure if the execution dies after claiming the key and before the write actually lands.

I prefer a small state machine for the idempotency record:

  • in_flight — this event is currently claimed
  • completed — the external side effect is confirmed
  • poison/review — outcome is uncertain or the payload is unsafe to retry blindly

Then on a retry:

  • completed → return 200 and do nothing
  • in_flight → wait/short-circuit unless the claim is stale
  • poison/review → verify the external system before replaying anything
  • missing → claim, validate, then perform the write

For multi-step writes, I’d also keep the provider’s stable event ID separate from n8n’s execution ID. A new execution ID on every retry is not an idempotency key.

The regression test I’d use is deliberately replaying the exact same webhook twice, then simulating a failure after the claim but before/after the external write. If both cases recover without either duplicating or silently dropping the action, the guard is doing its job.

The read-vs-write labels are a great operational habit too — especially once a workflow has several side-effect nodes.

@daniel_Samuel, one correction: Data Tables don’t currently support a UNIQUE constraint on a custom column, so that part won’t work. Keep the Postgres unique-key approach for the atomic claim.

Also, don’t clear the key just because the HubSpot node errors. A timeout can happen after HubSpot created the note. Clearing the key then allows the retry to create another one. That’s the case for the poison/review state described above, not automatic replay.

For the webhook response, set Respond to Using ‘Respond to Webhook’ Node and send 200 after the event ID and payload are durably stored. You can acknowledge before creating the note, but you then need a recovery process for stored events that haven’t completed.

The durable write before returning 200 is the key boundary here.

In my experience using Inistate, I treat the webhook as the start of a long-running operation rather than storing a simple seen/not-seen flag. n8n first creates or claims one Inistate record using the provider’s event ID. The record can move through Received, Processing, Completed, and Needs reconciliation.

Once that record exists, the workflow can return 200 and continue with the HubSpot write. If HubSpot times out, the operation moves to Needs reconciliation instead of deleting the key or retrying blindly. A recovery workflow or a person can check whether the note exists before choosing the next action.

The event ID prevents duplicate delivery from creating another operation. The operation state handles the harder case where the side effect may have happened but n8n did not receive confirmation.