Ambiguous Tool Outcomes Benchmark
Counterfactual benchmark for LLM agent recovery from ambiguous tool outcomes
Counterfactual benchmark for LLM agent recovery from ambiguous tool outcomes
Mutation testing of task-scoped state oracles in software-agent benchmarks
JSON Schema serialization-order robustness in black-box LLM generation
VEX V5 robot control code — served as the lead programmer, won 2nd Prize at the SJTU campus competition
Grammar notes compiled for the SPEIT French curriculum
A Ren’Py suspense visual novel — uncovering the truth behind a death in a snowbound school