AgentAtlas — Multi-Axis Benchmark for LLM Agent Behavior
arXiv researchers propose AgentAtlas, a framework that moves beyond single accuracy scores to evaluate LLM agents across six decision states and nine failure modes.
• Control-decision taxonomy: Act, Ask, Refuse, Stop, Confirm, Recover
• Trajectory-failure labels capture both root cause and downstream impact
• Addresses fragmentation across tool-call validity, safety, robustness, consistency
Reflects growing consensus that deployable agents need richer evaluation than outcome leaderboards.