Ambiguous Tool Outcomes Benchmark

A counterfactual benchmark studying how language model agents recover when a tool returns ambiguous outcomes. The benchmark evaluates whether agents can detect ambiguity, seek clarification, or recover to a correct task completion instead of proceeding on a wrong assumption.

Tech: Python

View on GitHub