Ax Corby Rosset, Pratyusha Sharma, Andrew Zhao, Miguel Gonzalez-Fernandez, Ahmed Awadallah 4/9/2026

The Art of Building Verifiers for Computer Use Agents

Framework for building verifiers for computer use agents with principles for constructing rubrics and training signals.

Ax Ahsan Bilal, Muhammad Ahmed Mohsin, Muhammad Umer, Asad Aali, Muhammad Usman Khanzada, Muhammad Usman Rafique, Zihao He, Emily Fox, Dean F. Hougen 4/9/2026

$S^3$: Stratified Scaling Search for Test-Time in Diffusion Language Models

S³: stratified scaling search for test-time inference in diffusion language models using verifier-guided sampling to improve generation quality without retraining.

Ax Hongyi Lu, Nian Liu, Shuai Wang, Fengwei Zhang 4/9/2026

ClawLess: A Security Model of AI Agents

ClawLess: security framework enforcing formally verified policies on autonomous LLM-based AI agents to mitigate code execution and data retrieval risks.

Ax Mingchen Zhuge, Changsheng Zhao, Haozhe Liu, Zijian Zhou, Shuming Liu, Wenyi Wang, Ernie Chang, Gael Le Lan, Junjie Fei, Wenxuan Zhang, Yasheng Sun, Zhipeng Cai, Zechun Liu, Yunyang Xiong, Yining Yang, Yuandong Tian, Yangyang Shi, Vikas Chandra, J\"urgen Schmidhuber 4/9/2026

Neural Computers

Proposes Neural Computers (NCs), a model architecture unifying computation, memory, and I/O as learned runtime state, aiming toward completely neural computing systems.

Ax Manish Bhatt, Sarthak Munshi, Vineeth Sai Narajala, Idan Habler, Ammar Al-Kahfah, Ken Huang, Blake Gatto 4/9/2026

The Defense Trilemma: Why Prompt Injection Defense Wrappers Fail?

Theoretical proof that no continuous wrapper defense can prevent all prompt injections in LLMs with connected prompt space, characterizing defense failure modes.