Tag: Agent Graph Execution

Sep 20
Step-Level Precision vs. Global Task Success: Micro- and Macro-Level Evals for Agent Graph Execution

When systems engineers begin evaluating autonomous artificial intelligence agents, they immediately confront an evaluation paradox: an agent can execute every intermediate tool call with apparent syntactic precision, yet fail completely to achieve the user’s high-level business goal. Conversely, an agent can stumble through clumsy, redundant, or malformed intermediate steps, trigger multiple retry warnings, and still […]