Ambiguous Tool Outcomes Benchmark
A counterfactual benchmark studying how language model agents recover when a tool returns ambiguous outcomes. The benchmark evaluates whether agents can detect ambiguity, seek clarification, or recover to a correct task completion instead of proceeding on a wrong assumption.
Tech: Python
