Intake record guard: page a human when an AI intake agent logs an emergency nobody can act on

I build AI intake agents (phone, SMS, web form) for home service contractors. During testing, my agent correctly told a caller with a gas smell to leave the house, then logged the emergency with a null address and a null callback number. The escalation was useless: nobody could be sent and nobody could call back.

This workflow sits between the agent and everything downstream:

  • Emergency with no usable callback number or address: pages a human immediately, naming what’s missing
  • Any other emergency: pages a human with the address and callback number
  • Missing or placeholder fields (“unknown”, “123 Main St”, “555-555-5555”): flagged for review
  • Everything: continues to your CRM or sheet

The webhook also replies with severity (page, review or ok) and the list of problems, so the voice or SMS platform can log the result.

Nodes: Webhook → Code → Respond to Webhook → IF → HTTP Request (Slack) and a NoOp placeholder for your CRM. No credentials included. Swap the Slack URL for Discord or anything that accepts a POST.

Repo with test commands (PowerShell and curl): GitHub - rizkynandapr/n8n-intake-record-guard: n8n workflow that pages a human when an AI intake agent logs an emergency with no address or callback number. · GitHub
The story behind it: 8 of my AI agent's 30 test calls failed. Every one was my fault. - DEV Community

Feedback welcome, especially on the placeholder list: which fake values have your agents invented to fill a required field?

Love this pattern. The guard between the agent and the downstream systems is the part most people skip, then one bad escalation ruins trust in the whole thing.

On placeholder values, here is what I have seen agents invent to fill required fields: N/A, none, unknown in both cases, TBD, test, John Doe, asdf, phone numbers like 555-555-5555 and 123-456-7890, 999 Main Street, 1 Main St, 0, single characters like x or a dash, and dates like 01/01/2000 or today’s date when a date was required but unknown. The sneakiest one is a real-looking name and number the model hallucinated from the conversation context, which passes a naive null check.

One addition worth considering: check the callback number against the number the caller is actually calling from. If the agent logged a different number than the caller ID, that is a strong signal it invented one.

Thanks, this is exactly the list I was hoping for. Several of yours aren’t in mine yet: 123-456-7890, asdf, single characters like x or a dash, 0, 1 Main St, 999 Main Street, and the 01/01/2000 date. I’ll add them to the Code node.

The caller ID check is a better signal than anything on the list, because it catches the hallucinated real-looking number that passes every placeholder check. The record alone doesn’t carry the caller ID, so I’ll add it as an optional check that runs when the phone platform sends the number alongside the record.

Done: both are in the repo now. The placeholder list includes yours, callback numbers need at least ten digits and can’t sit in 555-0100–0199 (the range reserved for fiction), and there’s an optional caller ID check. Send caller_id next to the record and a mismatch gets flagged for review. Credited you in the README.

Really useful build. The gas-leak-with-no-address story is the perfect example of why a guard between the agent and downstream systems matters, and the caller ID check is a great addition.

One thing I’d look at next is what happens after the page. Right now the guard catches the bad record, fires a Slack message, and the execution ends. But an emergency page that nobody acknowledges is the same silent failure as before, just moved one step down the line. The agent did its job, the guard did its job, and still nobody called the customer back.

The page isn’t the end of the workflow. It’s the start of a long-running job. It might take 2 minutes or 40 for someone to act, and that waiting doesn’t belong inside one execution. I’d give each flagged record its own status and let it move through a small state machine:

Paged → Acknowledged → Callback Made → Resolved
plus Escalated (no ack within N minutes)

That gives you a few things the fire-and-forget page can’t:

  • Escalation on silence. A scheduled check finds records still in Paged after, say, 5 minutes and pages the next person or calls the owner. No ack is a signal in itself.
  • review records stop instead of flowing on. At the moment every record continues to the CRM, including the ones with “123 Main St”. If review is a state, the record waits there until a human fixes the address or callback, then moves to the CRM clean. Bad data doesn’t land in the CRM with a flag nobody reads.
  • Proof someone acted. Each transition records who and when. For a gas emergency, “paged at 14:02, acked by Dave at 14:04, callback at 14:09” is exactly what you want to be able to show a customer, or an insurer, afterwards.

The guard stays exactly as you built it. It decides page / review / ok. The only change is that the decision creates a record with a status, instead of a message that disappears into a Slack channel.

The review case is the one that stings. You’re right that a flagged record still lands in the CRM with a flag, and a flag nobody reads is the same as no flag. That’s a hole in what I shipped, not a nice-to-have.

The escalation-on-silence part I want to be careful about. A state machine needs somewhere to keep the state, and the moment the guard needs a table it stops being a file you import and activate in two minutes. That’s the trade I’d be making for everyone who just wants the page to fire.

So I’d split it. The guard keeps deciding page / review / ok and stays import-and-go. A second workflow, optional, owns the states: it writes the row, runs a schedule that finds anything still in Paged after N minutes, and holds review rows until a human clears them. Anyone already running Airtable, Postgres or n8n’s own data tables points it at that.

One thing I don’t know and would rather ask than guess: who acknowledges, and how? A Slack button is the obvious answer but it needs a public callback URL, which plenty of self-hosted setups don’t have. A reply in the thread is uglier and works everywhere. What have you seen hold up with on-call people who are out driving a van?

The “paged at 14:02, acked by Dave at 14:04” line is the part I hadn’t thought about at all. That isn’t monitoring, it’s evidence, and for a gas call it’s what you’d actually want six months later.

The split makes sense. The guard stays import-and-go, and the state workflow is opt-in for anyone who needs it.

On who acknowledges and how, one rule covers most of it: anything the provider pushes to you needs a public URL, and anything you poll doesn’t.

  • Slack, no buttons: a :white_check_mark: reaction or a “mine” reply in the thread, with n8n polling the message every minute. A reaction is one tap from a phone notification.
  • Telegram with buttons: inline buttons work without a public URL if you poll getUpdates on a schedule instead of using the webhook trigger. For people out driving, one big button on the lock screen holds up best.
  • Escalate on silence: no ack after 5 minutes, page the next person, then call.

For the state side, I use Inistate as the record layer. Each page is a record in Paged / Acknowledged / Called Back / Resolved. The on-call person acks from the app on their phone, and n8n only reads and moves the record, so there’s no callback URL to expose. The “paged 14:02, acked by Dave 14:04” trail is kept automatically as the record’s history, so the evidence comes for free. A data table or Postgres works too, you just build that log yourself.

One more state worth adding: Called Back. “Acked but never called the customer” is the next silent failure, and for a gas call it’s the one that matters six months later.

The push/poll rule is the part I was missing. I kept reaching for buttons,
hitting the callback URL wall, and stopping there. It never occurred to me
that polling turns the same button into something a self-hosted box can use.

One catch for anyone copying the Telegram trick: getUpdates and a set
webhook can’t coexist. If you’ve ever pointed Telegram at an n8n webhook
trigger, getUpdates returns 409 until you call deleteWebhook. Worth saying
out loud because the error doesn’t explain itself.

Called Back is the right call and I’m annoyed I didn’t see it, because it’s
the same shape as the bug that started all this. The original failure was an
escalation nobody could act on. “Acked but never called” is an escalation
someone said they’d handle and then didn’t, which is worse, because now
there’s a record saying it was handled.

Which raises the one I don’t have an answer for: what happens when they do
call back and the customer doesn’t pick up? That isn’t Resolved and it isn’t
really Called Back either. My guess is it needs an attempt counter and a rule
like two attempts in fifteen minutes, then back to Paged for someone else. Do
you keep that as a state or as a field on the record?

On storage I’ll keep the workflow agnostic, reading and writing through one
node people can repoint, since whatever you already run is the right answer
for you. I’ll look at Inistate, I hadn’t come across it.

Mapping I’m working to now: Paged → Acknowledged → Called Back → Resolved,
plus Escalated on no ack after 5 minutes.

Good catch on the 409. That one costs people an hour, because nothing in the error points at the old webhook. Worth putting in the README next to the Telegram setup.

On the no-answer case, I’d use both, and the rule I use to split them is: if it changes what’s allowed to happen next, it’s a state. If it’s just counting, it’s a field.

A failed callback changes what’s allowed next. It can’t be Resolved, it needs a retry timer, and someone still owns it. So I’d make it a state:

Acknowledged → Callback Attempted → Called Back → Resolved

Callback Attempted loops back to itself on each try, while the attempt count and the time of each try live as fields (or as entries in the log). The rule then reads naturally: “in Callback Attempted with 2 or more attempts in 15 minutes → move on.”

The part I’d think hard about is where it moves on to. Sending it back to Paged for someone else fits when the first person can’t follow through. But if the customer isn’t picking up, the next person rings the same phone. That’s a different problem, so I’d give it its own exit:

Callback Attempted → Customer Unreachable, owned by a supervisor rather than the on-call pool. For a gas call, “we can’t reach the person who reported it” is a decision for a person to make, not a timer: send someone, try a second number, or follow your emergency guidance. Sending an SMS on the first failed attempt also helps, since some people won’t answer unknown numbers but will call back after a text.

Your mapping looks right to me. The one change I’d make is to keep Escalated pointed back into Acknowledged, so a page that escalated and was then acked follows the same Called Back path. Otherwise escalated pages turn into the place where the trail ends.

The state-versus-field rule is the part I’m keeping. “If it changes what’s allowed to happen next” settles an argument I’ve been having with myself for a week, and it generalises well past this workflow.

Customer Unreachable owned by a supervisor is right, and for the reason you gave rather than the obvious one. I’d have sent it back to the pool, and the next person would have rung the same dead phone. That looks like progress on the board and is nothing.

One thing I’d layer on top, from the case that started this. What you’re allowed to do in Customer Unreachable depends on what the record actually holds:

address, no usable phone -> you can still send someone. For a gas call
that's the whole answer.

phone, no address -> you can't send anyone, so the retry ladder is all
you have, and the SMS matters more.

neither -> this is the record that started the thread. Nothing to retry,
nothing to send to, and the only thing left is whatever the agent
captured in the transcript.

So Customer Unreachable probably wants a reason on it, or separate exits. I’ll build it as one state with a reason field and see which one hurts first.

I’d also shorten the ladder by emergency type. Two attempts in fifteen minutes is fine for a burst pipe. For a gas call I’d go to the supervisor on the first failed callback, because those fifteen minutes are being spent by someone standing in the house.

Escalated pointing back into Acknowledged: agreed, no argument. An escalated page that then goes quiet is exactly the hole I’d have left open.

The 409 is in the README now, next to the rest of the acknowledgement notes. Thanks for staying on this one.

1 Like

One gap sits under the whole state machine: every state (Paged, Acknowledged, Called Back, Customer Unreachable) rests on a record written by the same agent that made the original error. When the gas-smell call logged null address and null callback, nothing can tell you later whether the agent logged those nulls at call time or something dropped them downstream. Months later, “paged at 14:02” proves a page fired, but nothing proves what the agent actually knew.

The fix that generalizes: capture the agent’s full call as one tamper-evident unit: inputs, tool invoked, exact returned result. If the address comes back null, the null is inside the record with a hash over it, so nobody can alter it silently later. A dropped field becomes tamper-evident fact, not an untrustworthy log line.

This is the pattern Zambo receipts implement: a verifiable receipt is proof the saved result was not changed after execution. You can see one at Check a receipt | Zambo.

Concrete move: have the voice platform, not the agent, snapshot the tool result, and page off that snapshot instead of the agent’s log.

The first part is fair and I’d separate it from the rest. The record is
written by the same agent that made the mistake, so treating it as evidence
of what the agent knew is circular. That’s the reasoning behind the caller ID
check in this workflow: it reads a value the agent never sees, which makes it
the only field in the chain that doesn’t come from the thing being checked.

Your concrete move follows from that and costs nothing. Have the platform
snapshot the tool result and page off the snapshot rather than the agent’s
log. Worth doing whether or not anything is signed.

Where I don’t follow is the tamper-evident part. Nothing in this failure
involved anyone altering a record. The gas call wrote nulls at write time and
those nulls were honest. A hash over that record gives me a provably
unaltered wrong record, which is the same problem with a receipt stapled to
it. What I need is a second source that disagrees with the agent, and a
signature over a single source can’t be that.

If you have real cases where a correct record was silently changed
downstream, I’d read them, because that’s a failure mode I haven’t met and it
would change what I build. Short of that I’d put the effort into independent
capture rather than into proving the agent’s own output wasn’t edited.