Delegation needs oversight.
As agents take on longer workflows, users move from making each
decision themselves to overseeing execution. They still need to judge
whether the agent’s choices serve their goals. Even reasonable choices
can change task outcomes or how results should be interpreted, making
them worth user verification.
Yet following every action becomes difficult as workflows grow.
Evidence is scattered across messages, code, tool results, and
intermediate artifacts, leaving users to reconstruct what happened
and why it matters. Effective oversight therefore requires identifying
consequential decisions and connecting them to the evidence users
need to assess their implications.
AgentMonBench evaluates these needs in
software-engineering tasks along two complementary dimensions:
alignment between requirements and behavior, and awareness and
verification of consequential decisions. To support these judgments,
the
Evidence-Grounded Behavior Graph (EBG) organizes
source-linked evidence around observed behaviors and their
relationships, helping monitors connect the agent’s choices to user
requirements and potential consequences.