This is the exact distinction that led me to build UAEP:
EXECUTION SUCCESS ≠ VERIFIED DESTINATION REALITY
A clean trace establishes what the workflow recorded. It does not necessarily establish that the authorised destination objective became true.
I would also be careful about treating a returned transaction_id or message_id as automatic proof. An authentic identifier can still belong to the wrong resource, version, execution or environment. A provider can accept an operation without the final destination satisfying the objective.
UAEP separates:
• the agent or workflow claim;
• the executor or tool result;
• independently observed destination reality;
• whether the evidence applies to the exact objective, execution, resource and authority;
• whether recovery actually closes and reverifies the original objective.
The result becomes VERIFIED only when the required destination facts are established. If an action may have committed but its response was lost—or the evidence is stale, incomplete or mismatched—the truthful result may remain UNKNOWN / NOT_VERIFIED rather than being forced into SUCCESS or FAILURE.
UAEP is not just a theory. Its current implementation is complete and internally qualified within its declared evaluation boundary. It has been exercised against false executor success, wrong destinations, stale and wrong-resource evidence, partial effects, ambiguous commits, unsafe retries, conflicting observations and recovery that ran without closing the original objective.
Within those bounded qualified demonstrations:
FALSE VERIFIED = 0
That is bounded evidence, not a claim that UAEP cannot fail. I have opened the implementation to a controlled falsification challenge:
For an n8n workflow, the key design question is:
What destination-connected observer can establish the required state independently of both the AI Agent node and the tool path that produced the success claim?