Ambiguous Tool Outcomes Benchmark
Counterfactual benchmark for LLM agent recovery from ambiguous tool outcomes
Counterfactual benchmark for LLM agent recovery from ambiguous tool outcomes
Grammar notes compiled for the SPEIT French curriculum
JSON Schema serialization-order robustness in black-box LLM generation
Mutation testing of task-scoped state oracles in software-agent benchmarks
A Ren’Py suspense visual novel — uncovering the truth behind a death in a snowbound school
VEX V5 robot control code — served as the lead programmer, won 2nd Prize at the SJTU campus competition
Published in Research Square (preprint), 2026
An empirical study of how validation-equivalent JSON Schema serialization order affects black-box LLM generation, with practical guidelines for schema design.
Recommended citation: Sun, S. (2026). "Testing JSON Schema Instruction Artifacts: Distributional Robustness under Validation-Equivalent Serialization and JSON Mode." Research Square. doi:10.21203/rs.3.rs-10610100/v1
Download Paper
Published in Research Square (preprint), 2026
A cross-benchmark empirical study applying mutation testing to task-scoped state oracles to calibrate the reliability of software-agent benchmarks.
Recommended citation: Sun, S. (2026). "Mutation Testing of Task-Scoped State Oracles in Software-Agent Benchmarks: A Cross-Benchmark Empirical Study." Research Square. doi:10.21203/rs.3.rs-10665114/v1
Download Paper
Published in Research Square (preprint), 2026
A counterfactual evaluation framework studying whether LLM agents can detect, clarify, and recover from ambiguous tool outcomes instead of proceeding on wrong assumptions.
Recommended citation: Sun, S. (2026). "Did It Happen? Counterfactual Evaluation of LLM Agent Recovery from Ambiguous Tool Outcomes." Research Square. doi:10.21203/rs.3.rs-10730245/v1
Download Paper