Anchor: fixing mismatched task specs in agent benchmarks
arXiv paper tackles “artifact drift”—when task instructions, environments, and reward specs disagree, leaving benchmarks unsolvable or gameable.
Anchor converts domain expert workflows into constraint optimization programs, then jointly generates consistent instructions, environments, and verifiers from a single spec.
Targets enterprise agent evaluation: long-horizon business tasks that demand realism, verifiability, and scale at once.