Publications

Preprints | 预印本

Did It Happen? Counterfactual Evaluation of LLM Agent Recovery from Ambiguous Tool Outcomes

Published in Research Square (preprint), 2026

A counterfactual evaluation framework studying whether LLM agents can detect, clarify, and recover from ambiguous tool outcomes instead of proceeding on wrong assumptions.

Recommended citation: Sun, S. (2026). "Did It Happen? Counterfactual Evaluation of LLM Agent Recovery from Ambiguous Tool Outcomes." Research Square. doi:10.21203/rs.3.rs-10730245/v1
Download Paper

Mutation Testing of Task-Scoped State Oracles in Software-Agent Benchmarks: A Cross-Benchmark Empirical Study

Published in Research Square (preprint), 2026

A cross-benchmark empirical study applying mutation testing to task-scoped state oracles to calibrate the reliability of software-agent benchmarks.

Recommended citation: Sun, S. (2026). "Mutation Testing of Task-Scoped State Oracles in Software-Agent Benchmarks: A Cross-Benchmark Empirical Study." Research Square. doi:10.21203/rs.3.rs-10665114/v1
Download Paper

Testing JSON Schema Instruction Artifacts: Distributional Robustness under Validation-Equivalent Serialization and JSON Mode

Published in Research Square (preprint), 2026

An empirical study of how validation-equivalent JSON Schema serialization order affects black-box LLM generation, with practical guidelines for schema design.

Recommended citation: Sun, S. (2026). "Testing JSON Schema Instruction Artifacts: Distributional Robustness under Validation-Equivalent Serialization and JSON Mode." Research Square. doi:10.21203/rs.3.rs-10610100/v1
Download Paper