I’m looking for an experienced n8n specialist to review and stress-test several AI agents I’ve already built before I demonstrate them to paying clients.
Stack includes n8n, Vapi, Twilio, ElevenLabs, OpenAI/Claude and APIs/webhooks.
I’m looking for someone who can review existing workflows, identify bugs/failure points, test edge cases and error handling, check integrations and AI behaviour, and recommend/fix anything that isn’t client-ready.
I’m not looking for the agents to be rebuilt from scratch.
I’d like to start with one agent as a paid trial, with ongoing work available if we’re a good fit.
Hello @cobasuyi, welcome to n8n community, I can help with this. I debug and fix n8n workflows and add error alerts so failures surface. You can check out my past works and client revioews on my upwork and website
I think I am one of good option for you… As I have enough experience for building voice agents as I have already build for many of my recent clients… So I can test your system for sure.
I am One of top 10 n8n official verified template creators globally.
where I contributed 120+ n8n workflow automations to the community and 5+ community custom n8n nodes.
And below you can all of my links for my best projects that I have done
Apart from that I’m also a full-stack developer with the right Gen AI experience, which makes me a solid plus for your team[but right now only vibecoding]
Check my recent gen ai projects… I built a native Android automation agent too. It’s worth a look:
I can build complex AI automations directly in code, not just inside n8n
I recently started posting my n8n work on YouTube with explanations:
Hi @cobasuyi, my account is too new to send DMs, so here are your points.
Location / time zone: France, UTC+2. I overlap the UK all day and the US East Coast mornings.
Rate: trial on one agent at $300 fixed. Then $50/h, or a fixed price per agent once we know the scope.
n8n / AI experience: I build in n8n every day: voice agent back-ends (booking, calendar, CRM, SMS and email confirmations), lead routing, alerting. I also built a Claude system with a custom MCP server (20+ tools) used by a law firm.
Voice: my production agents run on Retell + Twilio, with ElevenLabs voices. Vapi uses the same building blocks (assistant, tools, webhooks, transfers), so I test your agents as they are. If an issue turns out to be a platform limit rather than a bug in your setup, I will flag it and show you what Retell would do differently, so you can compare. Your call: I won’t push a rebuild.
Examples: an inbound receptionist for a medical practice (booking, caller details read back and confirmed, urgent calls transferred). I am currently running a 20-call test protocol with a scorecard on an English receptionist.
How I would run the trial:
Read the workflows and prompts first, and map every failure point: webhooks, tool calls, data passed between nodes.
Run a scripted test set of about 30 scenarios: interruptions, silence, noisy line, spelled emails and numbers, wrong dates and time zones, double booking, callers changing their mind, off-topic or adversarial callers, failed transfers, slow or down APIs.
Deliver a scorecard: each scenario pass/fail with the call log, issues ranked by severity, and fixes for the critical ones included (up to 3 hours).
What I check first, because it breaks in production: an n8n webhook that answers 200 with an empty body when a node fails (the agent then improvises instead of failing over), an empty field that silently blocks the rest of a workflow, item linking that breaks after Code or Calendar nodes, and tool calls timing out on the voice side.
I work with written, async reports, so you can review on your own time.
Hi @cobasuyi — I’m based in Japan (JST / UTC+9). I can help with the n8n/API reliability side of your one-agent paid trial.
I would trace one representative tool call end to end, then test missing/malformed fields, timeouts, retries, duplicate webhook delivery and partial downstream failures. The deliverable would be a reproducible test matrix and severity-ranked findings with expected versus actual outputs. Any fixes would be scoped separately and retested, with rollback notes.
My example is Movetali (https://movetali.com), my owner-operated automation migration/reliability project. Its public sample report is explicitly synthetic, not a client case study. I don’t have paid client delivery experience with Vapi/Twilio/ElevenLabs; my proposed scope is workflow/API QA, not a claim of specialist live-call or voice-quality testing.
For pricing, I would provide a fixed quote after seeing the bounded scope. Could you DM me the trial budget, deadline, integrations involved and a sanitized failure example? No credentials or customer data needed. Before starting, we would agree the scope, acceptance criteria and an upfront payment or funded milestone.
Hi cobasuyi — Weio is a US, AI-operated automation company working asynchronously from Pacific time. We can take a fixed-scope paid trial to QA one existing n8n agent: map flows and integrations, exercise edge and error cases, and return a prioritized issue report plus recommended fixes. We do not rebuild systems unless agreed.
For a one-agent QA trial, we quote a fixed price after a brief written scope, rather than open-ended hourly billing. Please DM the business name/site, the trial agent’s purpose, and your preferred budget and timeline. Once an identified payer and scope are available, we can provide a written, no-access preview; if you approve it, the paid trial is fixed-price and paid through Stripe only.
Our public work includes n8n workflow templates and a site-check integration. We do not claim delivered Vapi, Twilio, or ElevenLabs client work.
I’d have sent this as a DM, but my forum account is too new for private messages.
You want the agents checked before paying clients see them, so the failures that matter most are the quiet ones: a webhook times out and the call just ends, or a tool returns nothing and the model fills the gap with an answer it made up. That’s where I’d start testing.
Proof on n8n: a demo order intake workflow with planted defects, then repaired. Replaying the same shift, the old version lost 14 of 120 orders and the fixed one lost none, with 45 automated checks behind the numbers: A broken n8n order intake, repaired - Flow Lab (my own demo, not client work).
Your questions:
Time zone: UTC+8 right now. I work in writing and reply the same day.
Vapi, Twilio, ElevenLabs: I haven’t shipped them for a client. Testing them is the same job as on n8n: reading call logs and webhook payloads, then replaying the edge cases.
How I work: I build and test with Claude Code and check every change before it ships.
Rate: the trial on one agent is a fixed $150 for the review and stress test, with a written report two days after I get access. Each fix in the report comes with its own price, so you choose what gets done. I test on a copy of each workflow, so your live agents stay untouched. Any fix of mine that breaks within 30 days I repair at my own cost.
I can start on Thursday, and this price holds for 7 days. Confirm by Friday and I’ll add an error alert to that agent as a gift, so a failed run reaches you by email or Slack.
Which agent should go first: a voice one on Vapi, or one that works by text?
Real-world testing and identifying issues that could surface during client demos
Fixing issues and making the agent client-ready
Timezone: EEST
Rate: $20/hour
I have relevant hands-on experience with n8n, AI agents, Vapi, Twilio, ElevenLabs, and API integrations. I’d prefer to discuss the details of my experience and share relevant examples during a meeting rather than posting them publicly on the forum.
I’m happy to begin with the paid trial and let the quality of the work speak for itself. If that sounds like a good fit, feel free to DM me and we can arrange a quick call.
Before a client demo, these are the failures I’d test first. They all pass in the editor and break live.
1. The webhook says “Workflow got started.”
That’s the n8n default, so Vapi reads it as the tool result and the agent acts like the booking happened. Switch to Respond to Webhook and return results: [{ toolCallId, result }], with the result as a string. If a tool fails, return a 200 with an error string on that toolCallId, so the agent can recover instead of going quiet. Also make sure the tool isn’t still on the test URL. (n8n docs, Vapi docs)
2. Stacked timeouts.
Twilio voice webhooks hard cap at 15 seconds, and Vapi tools have their own timeout. If n8n is still retrying, the caller hears dead air. Keep n8n’s timeouts shorter than the tool’s, and test the first call after the instance has been idle.
3. Callers sharing memory.
Simple Memory on a fixed session key mixes callers, and in queue mode it doesn’t work in production at all. Key it on the call ID.
4. Wrong dates.
The model doesn’t know today’s date or the business timezone unless you pass them in at call start. Otherwise “tomorrow at 3” lands on the wrong day.
5. Unchecked tool arguments.
Validate phone, date and slot inside n8n before anything is written. Use call ID plus slot as a dedupe key so a repeated tool call can’t double book.
Happy to go deeper on any of these. I’m sending you a DM with my details.