The problem
Automated evaluation is assumed reliable for tool-using language agents because human annotation does not scale. That assumption had not been audited against human labels at scale.
04Research artifact · arXiv:2604.16706
Code and data for “Auditing Automated Evaluation, Error Propagation, and Runtime Mitigation in Tool-Using Language Agents.” It audits the widely-assumed reliability of automated evaluation for tool-using LLM agents against human annotation.
Sole author · v1 April 2026 · revised 9 August 2026 · arXiv preprint · 11 pages, 4 figures, 8 tables
Automated evaluation is assumed reliable for tool-using language agents because human annotation does not scale. That assumption had not been audited against human labels at scale.
Substring-heuristic judging — the default in much of the agent-evaluation literature — is inadequate. The paper audits automated evaluation against human annotation, traces how errors propagate through multi-step agent execution, and examines what runtime mitigation can actually recover.
Code and data are public at github.com/bhaskargurram-ai/agenthallu-bench. 11 pages, 4 figures, 8 tables; v1 April 2026, revised 9 August 2026.