¿Tu agente alguna vez ha afirmado haber hecho algo que en realidad no sucedió?

No es un error. No es un fallo. El agente dice “correo enviado” o “reembolso procesado”, el rastro se ve limpio, sin excepciones en ningún lado — y la cosa nunca sucedió. Después de ocurrir esto varias veces, me molesta que todas las herramientas de observabilidad que he probado lo reporten como un éxito, porque desde el punto de vista del rastro sí lo es.

Dos preguntas para quien esté ejecutando agentes o flujos de trabajo en producción:

¿Te ha pasado esto?

Si es así — ¿cómo lo descubriste? ¿Por una queja del cliente, o fue algo que lo detectó?

Tengo curiosidad genuina por saber si es algo común o si simplemente he construido las cosas mal.

Hola Aditya,

Este es un problema común conocido como Silent Agent Failure (Fallo Silencioso del Agente). Ocurre cuando el LLM asume que una herramienta tuvo éxito basándose en la salida de texto, o cuando una herramienta falla silenciosamente sin generar un código de error.

Aquí se explica cómo prevenirlo y detectarlo:

  1. Verificación Determinística: Agrega un paso justo después del agente para verificar independientemente el resultado a través de API/Base de Datos (por ejemplo, verifica Stripe para asegurar que el reembolso esté realmente registrado).

  2. Requiere Pruebas de Ejecución: Obliga a las herramientas a devolver JSON estricto con IDs específicos (transaction_id, message_id). Si falta el ID, trata el paso como un fracaso.

  3. Errores de Subflujos Estrictos: Asegúrate de que los subflujos de las herramientas generen errores explícitos (On Error: Stop Workflow) en respuestas que no sean 200, para que los fallos no se oculten detrás de texto limpio.

  4. Trabajos Cron de Auditoría: Ejecuta verificaciones periódicas para que coincidan las acciones solicitadas con los registros del backend actual para detectar discrepancias temprano.

Gracias — el punto de verificación determinista es en el que yo también sigo llegando.

Volviendo a mi pregunta original: ¿has tenido uno de estos en producción tú mismo? Me interesa menos los patrones de prevención que cómo se veía cuando sucedió — qué reportó la ejecución y cuánto tiempo pasó antes de que alguien lo notara.

Pregunto porque todos con los que hablo están de acuerdo en que es un problema en principio, y casi nadie tiene la historia.

¡Buena pregunta! Para ser transparente, personalmente no he perdido dinero por una falla silenciosa de agente en mi propio entorno de producción en vivo todavía, principalmente porque nos advirtieron temprano que configuráramos pasos de verificación.

Sin embargo, un colega en el espacio de automatización experimentó exactamente este escenario con un agente de soporte que emitía reembolsos:

  • La configuración: El agente activó una herramienta para llamar a Stripe para reembolsos parciales.

  • Lo que reportó la ejecución: El registro de ejecución de n8n estaba completamente verde (200 OK), y el LLM respondió con confianza: “Emití un reembolso de $20.”

  • Lo que realmente sucedió: El subproceso encontró un error de código de moneda inválido, pero la herramienta de API capturó el error correctamente y devolvió { "success": false } dentro de un estado 200 OK. El LLM vio la respuesta de texto, ignoró la bandera false, y alucinó que la acción fue exitosa.

  • ¿Cuánto tiempo hasta que se notó? 3 días después, cuando el cliente hizo seguimiento preguntando dónde estaba su dinero.

Así que aunque es una historia simulada basada en patrones de la industria, ¡el problema de enmascaramiento de errores 200 OK en n8n es muy real!

Yes, it’s definitely happened to me, and I don’t think it’s a sign you’ve built things badly. The trace only knows the steps ran. It has no idea whether the thing you wanted actually happened in the real world.

The first time, I found out from a user. Since then I’ve done two things:

  1. Check the result after every important action. For example, look up the refund or email by the ID the tool returned. If it isn’t there, throw an error with a Stop and Error node.
  2. Run a daily reconciliation workflow. It compares what the agent says it did with what the external system shows, and alerts on anything that doesn’t match.

Also worth checking: in some cases the agent never called the tool at all and just said it did. The execution log shows this right away.

Curious whether yours were mostly the agent hallucinating the action, or the API accepting the request and failing later?

This is the exact distinction that led me to build UAEP:

EXECUTION SUCCESS ≠ VERIFIED DESTINATION REALITY

A clean trace establishes what the workflow recorded. It does not necessarily establish that the authorised destination objective became true.

I would also be careful about treating a returned transaction_id or message_id as automatic proof. An authentic identifier can still belong to the wrong resource, version, execution or environment. A provider can accept an operation without the final destination satisfying the objective.

UAEP separates:

• the agent or workflow claim;
• the executor or tool result;
• independently observed destination reality;
• whether the evidence applies to the exact objective, execution, resource and authority;
• whether recovery actually closes and reverifies the original objective.

The result becomes VERIFIED only when the required destination facts are established. If an action may have committed but its response was lost—or the evidence is stale, incomplete or mismatched—the truthful result may remain UNKNOWN / NOT_VERIFIED rather than being forced into SUCCESS or FAILURE.

UAEP is not just a theory. Its current implementation is complete and internally qualified within its declared evaluation boundary. It has been exercised against false executor success, wrong destinations, stale and wrong-resource evidence, partial effects, ambiguous commits, unsafe retries, conflicting observations and recovery that ran without closing the original objective.

Within those bounded qualified demonstrations:

FALSE VERIFIED = 0

That is bounded evidence, not a claim that UAEP cannot fail. I have opened the implementation to a controlled falsification challenge:

For an n8n workflow, the key design question is:

What destination-connected observer can establish the required state independently of both the AI Agent node and the tool path that produced the success claim?