Ax Zhouzhou Shen, Xueyu Hu, Xiyun Li, Tianqing Fang, Juncheng Li, Shengyu Zhang 2/18/2026

World-Model-Augmented Web Agents with Action Correction

WAC: LLM-based web agent with world model for reasoning about environment changes and predicting execution risks before action.

Ax Gabriele Conte, Alessio Mattiace, Gianni Carmosino, Potito Aghilar, Giovanni Servedio, Francesco Musicco, Vito Walter Anelli, Tommaso Di Noia, Francesco Maria Donini 2/18/2026

RUVA: Personalized Transparent On-Device Graph Reasoning

On-device graph-based reasoning system for personalized AI with transparency and accountability, addressing limitations of RAG-based black-box systems.

Ax David Puertolas Merenciano, Ekaterina Vasyagina, Raghav Dixit, Kevin Zhu, Ruizhe Li, Javier Ferrando, Maheep Chaudhary 2/18/2026

Weight space Detection of Backdoors in LoRA Adapters

Detection method for backdoor attacks in LoRA adapters by analyzing weight space without requiring test data or knowledge of trigger patterns.

Ax Skyler Hallinan, Thejas Venkatesh, Xiang Ren, Sai Praneeth Karimireddy, Ashwin Paranjape, Yuhao Zhang, Jack Hessel 2/18/2026

OpaqueToolsBench: Learning Nuances of Tool Behavior Through Interaction

Benchmark dataset and framework for evaluating LLM agents' ability to learn tool behavior and improve documentation for opaque real-world tools through interaction.

Ax Mason Nakamura, Abhinav Kumar, Saswat Das, Sahar Abdelnabi, Saaduddin Mahmud, Ferdinando Fioretto, Shlomo Zilberstein, Eugene Bagdasarian 2/18/2026

Colosseum: Auditing Collusion in Cooperative Multi-Agent Systems

Colosseum framework for auditing collusive behavior in multi-agent systems where LLM agents communicate and coordinate through free-form language.

Ax Atticus Wang, Iv\'an Arcuschin, Arthur Conmy 2/18/2026

Automatically Finding Reward Model Biases

Approach to automatically detect reward model biases in LLMs using iterative LLM proposals, addressing spurious attributes like length, hallucinations, and sycophancy.

Ax Chengzhi Hu, Jonas Dornbusch, David L\"udke, Stephan G\"unnemann, Leo Schwinn 2/18/2026

Closing the Distribution Gap in Adversarial Training for LLMs

Method to improve adversarial training for LLMs by addressing distribution gaps that leave models vulnerable to simple in-distribution adversarial examples like prompt rewriting.

Ax Arya Tschand, Chenyu Wang, Zishen Wan, Andrew Cheng, Ioana Cristescu, Kevin He, Howard Huang, Alexander Ingare, Akseli Kangaslahti, Sara Kangaslahti, Theo Lebryk, Hongjin Lin, Jeffrey Jian Ma, Alexandru Meterez, Clara Mohri, Depen Morwani, Sunny Qin, Roy Rinberg, Paula Rodriguez-Diaz, Alyssa Mia Taliotis, Pernille Undrum Fathi, Rosie Zhao, Todd Zhou, Vijay Janapa Reddi 2/18/2026

GenAI for Systems: Recurring Challenges and Design Principles from Software to Silicon

Cross-stack analysis of generative AI applications in computing systems design from code generation through hardware design space exploration to RTL synthesis.

Ax Dongxu Zhang, Zhichao Yang, Sepehr Janghorbani, Jun Han, Andrew Ressler II, Qian Qian, Gregory D. Lyng, Sanjit Singh Batra, Robert E. Tillman 2/18/2026

Fast and Effective On-policy Distillation from Reasoning Prefixes

On-policy distillation technique for training student LLMs using teacher supervision on token-level trajectories, improving generalization over off-policy methods with reduced computational cost.