What Did the Agent Actually Do?Evidence-Grounded Oversight for Long-Horizon Agents

Zhongxiang Sun1*Jiahao Yan1*Hongkang Zhao1*Haojie Ding2Boheng Zhang2Fan Yang2Xiao Zhang1†Jun Xu1

1 Gaoling School of Artificial Intelligence, Renmin University of China   2 Kuaishou Technology

* Equal contribution   † Corresponding author

Figure 1. User oversight in long-horizon workflows. An illustrative research example connects a data-exclusion decision to its implications and supporting evidence. Click the figure to enlarge.

Delegation needs oversight.

As agents take on longer workflows, users move from making each decision themselves to overseeing execution. They still need to judge whether the agent’s choices serve their goals. Even reasonable choices can change task outcomes or how results should be interpreted, making them worth user verification.

Yet following every action becomes difficult as workflows grow. Evidence is scattered across messages, code, tool results, and intermediate artifacts, leaving users to reconstruct what happened and why it matters. Effective oversight therefore requires identifying consequential decisions and connecting them to the evidence users need to assess their implications.

AgentMonBench evaluates these needs in software-engineering tasks along two complementary dimensions: alignment between requirements and behavior, and awareness and verification of consequential decisions. To support these judgments, the Evidence-Grounded Behavior Graph (EBG) organizes source-linked evidence around observed behaviors and their relationships, helping monitors connect the agent’s choices to user requirements and potential consequences.

Abstract

As agents take on long-horizon tasks, users shift from making individual decisions to overseeing autonomous execution. Yet the volume of agent activity and the fragmentation of supporting evidence make it difficult to determine which decisions warrant user verification. We study monitors that identify consequential decisions and locate evidence to help users assess their implications.

We introduce AgentMonBench, a software-engineering benchmark comprising three subsets that cover two complementary dimensions: alignment between requirements and behavior, and awareness of consequential autonomous decisions for verification. To support these judgments, we propose the Evidence-Grounded Behavior Graph (EBG), a training-free method that groups source-linked evidence into behaviors and organizes their relationships into a graph. EBG presents task-oriented views of this graph to help monitors interpret behavior in context.

Experiments across eight models show that EBG improves decision identification and evidence localization in most settings compared with direct access to the original context. Further experiments show that EBG's evidence-localization gains persist across input scales and hyperparameter settings, while real-world applications illustrate its practical value for human oversight.

AgentMonBench

Two complementary dimensions, three tasks, 100 instances each. From input-side requirement gaps to output-side behavior changes and consequential decisions during execution.

Alignment between requirements and behavior

Are important requirements missing, or does the implementation diverge from what the user specified?

SpecGAP Input-side gaps
Find consequential repository behaviors left unspecified in incomplete requirements.
SilentSwap Output-side deviations
Detect semantic changes from documented behavior, even when existing tests pass.

Awareness and verification of consequential decisions

Which autonomous choices need user judgment, and what evidence reveals their implications?

FeedbackTrace Decisions during execution
Surface consequential choices in real interactions before the user intervenes. Later feedback defines the target; the monitor sees only the preceding history.
300evaluation instances100 per task
484omitted requirementsSpecGAP
500behavior substitutionsSilentSwap
100pre-feedback trajectoriesFeedbackTrace
INPUT SCALE

Long Contexts, Scattered Evidence

High-context FeedbackTrace inputs reach a median of 86.1k tokens, compared with about 32k for the repository tasks.

Group medians · 33 / 33 / 34 instances per task

SPECGAP

What the Brief Leaves Unsaid

484 omitted requirements span seven categories, from boundary behavior and API contracts to data flow and evidence.

7 requirement types · 4.84 omissions per instance

SILENTSWAP

Tests Pass. Behavior Changes.

500 semantic substitutions cover eight mechanisms. Every instance contains five changes that preserve existing tests.

8 change mechanisms · 5 substitutions per instance

FEEDBACKTRACE

Before the User Intervenes

100 pre-feedback trajectories contain 14,681 interaction events, with the longest spanning 908 events.

User prompts, assistant paragraphs, and tool exchanges

Input sizes: paper Table B1. Counts: released evaluation data. Click any chart to enlarge.

Download chart data

Three Oversight Tasks

Identify the consequential decision. Find the evidence behind it.

From requirements to behavior to consequential choices. Three illustrative examples from the paper. Click to enlarge, or explore each task below.

THE MONITOR'S JOB

Explore this task ↗
ILLUSTRATIVE SCENARIO

Across both dimensions

Identify the decision + Locate supporting evidence + Support user verification

Inside the paper

Evidence-Grounded Behavior Graph

EBG turns source-linked evidence into behaviors, scopes, and relations. Task-oriented graph views put each decision back in context.

Training-freeSource-linkedTask-oriented
01

Gather the evidence

Preserve source locations from requirements, code, traces, and feedback.

02

Organize behaviors

Group evidence into behaviors and connect their scopes and relationships.

03

Retrieve for the task

Expose relevant graph views so the monitor can reason with supporting context.

FIGURE / EBG From scattered artifacts to evidence-grounded oversight.

Leaderboard

Explore the paper's main results across eight models. Compare EBG with direct context access and RepoGraph on each oversight task.

AGENTMONBENCH

Research leaderboard

Download results

Scroll horizontally to compare all metrics.

Main experimental results; select column headers to sort.

Source: Table 1 of the paper. Scores are shown on a 0–100 scale; higher is better. Highlights above are derived from paired EBG–Base comparisons. These are reported experimental results, not a live submission ranking. Download full results (CSV).

EBG in Practice

A Codex review harness puts EBG to work in five research tasks adapted from real agent failures.

CASE STUDY / VALIDATION DATA LEAKAGE

The score went up.
The comparison broke.

A candidate reached 75.00% accuracy, above the baseline's 71.67%. Evidence-grounded review traced 8 of 80 added training samples to validation data. The agent disclosed the leakage and retained the baseline.

Explore the five case studies
01
OBSERVED GAIN

71.67% → 75.00%

The aggregate score favors the candidate using an extra training pool.

02
TRACE THE DATA

8 samples came from validation data

Review follows the added samples back to their provenance.

03
REVISE THE DECISION

Disclose the leakage. Keep the baseline.

The observed gain does not establish leakage-free generalization.

Also studiedUnclear objectivesChanged randomizationUnverified API callsUnequal search budgets

Five paired, illustrative case studies; not a separate aggregate benchmark. See Table 2 and the review-harness appendix in the paper.

Use the Codex Review Harness

Bring evidence-grounded review into a working agent session. Start with one of the paper's five research cases.

Before you start · Python 3.11+, Git, Node.js/npm, and a Codex CLI with hook support. Use your Python environment for all commands below.

  1. Check Codex

    Install or update with npm install -g @openai/codex@latest, then sign in with codex login. The hook-support check below must succeed.

    codex --version
    codex exec --dangerously-bypass-hook-trust --help
    codex login status
  2. Install

    Clone the code and install the standalone harness in your Python environment. A virtual environment is recommended.

    git clone https://github.com/zhk-lab/EBG.git
    cd EBG
    python -m pip install -e "./harness[case-studies]"
    python -m codex_harness --help
  3. Prepare a case

    Create an isolated workspace with MCP tools, hooks, and review skills installed.

    python harness/case_studies/prepare.py --install --round r1 --case 01_ambiguity
  4. Run with review

    Launch Codex with the generated integration. Use a new --round for a fresh run, or --model to select an available model. If several Codex versions are installed, use --codex /path/to/codex to select the executable that passed the check above. The ambiguity case pauses to ask for a metric priority before selecting a result.

    python harness/case_studies/run.py 01_ambiguity --round r1 --model gpt-6-luna --effort low

Review the decision. Check the evidence. Explain the finding.

The harness prompts reviews of ambiguity, material adjustments, and results. Use ebg_review to open a checkpoint, ebg_evidence to inspect its sources, and ebg_record to record the finding. Supported, consequential issues must also be explained to the user.

Connect the harness to your own workspace

From the EBG repository, generate the integration files:

python -m codex_harness --state-dir outputs/harness/demo setup --output outputs/harness/integration

Merge the generated config.toml MCP entry into your workspace's .codex/config.toml, place the generated skills in .agents/skills/, and load the generated hooks.json in .codex/. Start Codex from that workspace with codex -c features.hooks=true and approve the project hooks when prompted. Setup generates files; the case runner above performs the integration for the example workspace.

See the harness documentation and runner configuration for the full integration.

Resources

Read the paper, explore the implementation,
or evaluate your own monitor on AgentMonBench.

EBG · Method overview