هل ادعى وكيلك في أي وقت أنه فعل شيئاً لم يحدث فعلاً؟

ليس خطأ. ليس تعطل. يقول الوكيل “تم إرسال البريد الإلكتروني” أو “تمت معالجة المبلغ المسترد”، والتتبع يبدو نظيفاً، لا استثناء في أي مكان — والشيء لم يحدث أبداً. إذا حدث هذا عدة مرات، فإنه يزعجني أن كل أداة مراقبة جربتها تبلغ عنها كنجاح، لأنه من وجهة نظر التتبع فهو كذلك بالفعل.

سؤالان لأي شخص يقوم بتشغيل وكلاء أو مسارات عمل في الإنتاج:

هل حدث هذا لك؟

إذا كان الجواب نعم — كيف عرفت؟ شكوى من العميل، أم أن شيئاً ما اكتشفه؟

أنا فضولي حقاً ما إذا كان هذا شائعاً أم أنني فقط بنيت الأشياء بشكل سيء.

مرحباً أديتيا،

هذه مشكلة شائعة معروفة باسم Silent Agent Failure (فشل الوكيل الصامت). تحدث عندما يفترض نموذج اللغة أن الأداة نجحت بناءً على مخرجات نصية، أو تفشل الأداة بصمت دون رفع رمز خطأ.

إليك كيفية منعها واكتشافها:

  1. التحقق الحتمي: أضف خطوة مباشرة بعد الوكيل للتحقق بشكل مستقل من النتيجة عبر API/Database (مثل التحقق من Stripe للتأكد من تسجيل المبلغ المستردّ فعلياً).

  2. طلب إثباتات التنفيذ: أجبر الأدوات على إرجاع JSON صارم مع معرّفات محددة (transaction_id، message_id). إذا كان المعرّف مفقوداً، اعتبر الخطوة فاشلة.

  3. أخطاء سير العمل الفرعي الصارم: تأكد من أن سير العمل الفرعي للأدوات يرفع أخطاء صريحة (On Error: Stop Workflow) في استجابات non-200 حتى لا تُخفى الأخطاء خلف نص نظيف.

  4. مهام التدقيق المجدولة: قم بفحوصات دورية لمطابقة الإجراءات المطلوبة مقابل سجلات الخادم الفعلية للإشارة إلى عدم التطابق مبكراً.

شكراً — نقطة التحقق الحتمية هي التي أستمر في الوصول إليها أيضاً.

لكن بالعودة إلى سؤالي الفعلي: هل واجهت أحد هذه الحالات في الإنتاج بنفسك؟ أنا أقل اهتماماً بأنماط الوقاية منها بما يتعلق بكيف بدا الحال عندما حدثت — ما الذي أبلغت عنه المحاكاة، وكم من الوقت مضى قبل أن يلاحظ أحد.

أسأل لأن كل من أتحدث معهم يتفقون على أنها مشكلة من حيث المبدأ، وتقريباً لا أحد لديه القصة.

سؤال عادل! لأكون شفافًا، لم أخسر شخصيًا أموالًا بسبب فشل وكيل صامت في إعدادي الإنتاجي المباشر حتى الآن—في الأساس لأننا حذرنا في وقت مبكر لإعداد خطوات التحقق.

ومع ذلك، واجه زميل في مجال الأتمتة هذا السيناريو بالضبط مع وكيل يصدر استردادات:

  • الإعداد: قام الوكيل بتشغيل أداة للاتصال بـ Stripe لاسترجاع جزئي.

  • ما أبلغت عنه التتبع: كان تتبع تنفيذ n8n أخضر تمامًا (200 OK)، وردّ النموذج اللغوي بثقة: “أصدرت استردادًا بقيمة 20 دولارًا.”

  • ما حدث بالفعل: واجهت سير العمل الفرعي خطأ في رمز العملة غير الصحيح، لكن أداة API اكتشفت الخطأ بعناية وأرجعت { "success": false } داخل حالة 200 OK. رأى النموذج اللغوي استجابة النص وتجاهل علم false، وهلوس بأن الإجراء نجح.

  • كم من الوقت حتى تم اكتشافه؟ بعد 3 أيام، عندما تابع العميل السؤال عن مكان أمواله.

إذن بينما هي قصة محاكاة بناءً على أنماط الصناعة، فإن مشكلة إخفاء الأخطاء 200 OK في n8n حقيقية جدًا!

Yes, it’s definitely happened to me, and I don’t think it’s a sign you’ve built things badly. The trace only knows the steps ran. It has no idea whether the thing you wanted actually happened in the real world.

The first time, I found out from a user. Since then I’ve done two things:

  1. Check the result after every important action. For example, look up the refund or email by the ID the tool returned. If it isn’t there, throw an error with a Stop and Error node.
  2. Run a daily reconciliation workflow. It compares what the agent says it did with what the external system shows, and alerts on anything that doesn’t match.

Also worth checking: in some cases the agent never called the tool at all and just said it did. The execution log shows this right away.

Curious whether yours were mostly the agent hallucinating the action, or the API accepting the request and failing later?

This is the exact distinction that led me to build UAEP:

EXECUTION SUCCESS ≠ VERIFIED DESTINATION REALITY

A clean trace establishes what the workflow recorded. It does not necessarily establish that the authorised destination objective became true.

I would also be careful about treating a returned transaction_id or message_id as automatic proof. An authentic identifier can still belong to the wrong resource, version, execution or environment. A provider can accept an operation without the final destination satisfying the objective.

UAEP separates:

• the agent or workflow claim;
• the executor or tool result;
• independently observed destination reality;
• whether the evidence applies to the exact objective, execution, resource and authority;
• whether recovery actually closes and reverifies the original objective.

The result becomes VERIFIED only when the required destination facts are established. If an action may have committed but its response was lost—or the evidence is stale, incomplete or mismatched—the truthful result may remain UNKNOWN / NOT_VERIFIED rather than being forced into SUCCESS or FAILURE.

UAEP is not just a theory. Its current implementation is complete and internally qualified within its declared evaluation boundary. It has been exercised against false executor success, wrong destinations, stale and wrong-resource evidence, partial effects, ambiguous commits, unsafe retries, conflicting observations and recovery that ran without closing the original objective.

Within those bounded qualified demonstrations:

FALSE VERIFIED = 0

That is bounded evidence, not a claim that UAEP cannot fail. I have opened the implementation to a controlled falsification challenge:

For an n8n workflow, the key design question is:

What destination-connected observer can establish the required state independently of both the AI Agent node and the tool path that produced the success claim?