Ax Juyi Lin, Arash Akbari, Yumei He, Lin Zhao, Haichao Zhang, Arman Akbari, Xingchen Xu, Zoe Y. Lu, Enfu Nan, Hokin Deng, Edmund Yeh, Sarah Ostadabbas, Yun Fu, Jennifer Dy, Pu Zhao, Yanzhi Wang 5/12/2026

PhyGround: Benchmarking Physical Reasoning in Generative World Models

Benchmark suite for evaluating physical reasoning and dynamics accuracy in generative video world models.

Ax Edward De Brouwer, Carl Edwards, Alexander Wu, Jenna Collier, Graham Heimberg, Xiner Li, Meena Subramaniam, Ehsan Hajiramezanali, David Richmond, Jan-Christian H\"utter, Sara Mostafavi, Gabriele Scalia 5/12/2026

AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents

Benchmark for evaluating LLMs and agents on virtual cell modeling tasks, testing in silico phenotypic screening and biological discovery prediction.

Ax Linus Heck, Filip Mac\'ak, Roman Andriushchenko, Milan \v{C}e\v{s}ka, Sebastian Junges 5/12/2026

Shields to Guarantee Probabilistic Safety in MDPs

Formal framework for shielding techniques in probabilistic Markov decision processes, extending safety guarantees for autonomous agents with acceptable failure probabilities.

Ax Mohammadreza Armandpour, Fatih Ilhan, David Harrison, Ajay Jaiswal, Duc N. M Hoang, Fartash Faghri, Yizhe Zhang, Minsik Cho, Mehrdad Farajtabar 5/12/2026

Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Why

Analysis of on-policy distillation for training reasoning models, investigating when teacher-student supervision helps or hurts performance on token-level tasks.

Ax Yaxin Du, Xiyuan Yang, Zhifan Zhou, Wanxu Liu, Zixing Lei, Zimeng Chen, Fenyi Liu, Haotian Wu, Yuzhu Cai, Zexi Liu, Xinyu Zhu, WenHao Wang, Linfeng Zhang, Chen Qian, Siheng Chen 5/12/2026

DataMaster: Towards Autonomous Data Engineering for Machine Learning

Study of autonomous data engineering for ML systems, automating dataset discovery, adaptation, and validation to reduce manual data engineering workflows.

Ax Roxana Geambasu (Google,Columbia University), Mariana Raykova (Google), Pierre Tholoniat (Google), Trishita Tiwari (Google), Lillian Tsai (Google), Wen Zhang (Google) 5/12/2026

Engineering Robustness into Personal Agents with the AI Workflow Store

Research on engineering robustness into AI agents by applying traditional software engineering processes like testing, adversarial evaluation, and staged deployment instead of on-the-fly synthesis.

Ax Keya Hu, Linlu Qiu, Yiyang Lu, Hanhong Zhao, Tianhong Li, Yoon Kim, Jacob Andreas, Kaiming He 5/12/2026

ELF: Embedded Language Flows

Continuous diffusion language models using minimal adaptation to match effectiveness of leading discrete-token language model approaches.

Ax Himanshu Gupta, Shreyas Verma, Ujjwala Anantheswaran, Kevin Scaria, Mihir Parmar, Swaroop Mishra, Chitta Baral 5/12/2026

Polymath: A Challenging Multi-modal Mathematical Reasoning Benchmark

PolyMATH benchmark with 5,000 images evaluating multimodal LLM visual comprehension and abstract reasoning across 10 cognitive challenge categories.

Ax Haorui Wang, Jeff Guo, Lingkai Kong, Rampi Ramprasad, Philippe Schwaller, Yuanqi Du, Chao Zhang 5/12/2026

LLM-Augmented Chemical Synthesis and Design Decision Programs

LLM-augmented retrosynthesis system for chemical synthesis and drug development combining ML and LLMs to navigate combinatorial pathway space.

Ax Byeongchan Lee, Jonghoon Lee, Dongyoung Kim, Jaehyung Kim, Kyungjoon Park, Dongjun Lee, Jinwoo Shin 5/12/2026

Efficient LLM Collaboration via Planning

Planning-based framework for efficient LLM collaboration combining large and small models to reduce inference costs while maintaining performance.

Ax Yuanyi Wang, Yanggan Gu, Yiming Zhang, Qi Zhou, Zhaoyi Yan, Congkai Xie, Xinyao Wang, Jianbo Yuan, Hongxia Yang 5/12/2026

Model Merging Scaling Laws in Large Language Models

Empirical scaling laws for language model merging showing power law relationship between model size, expert number, and merging performance.

Ax Minhui Zhu, Minyang Tian, Xiaocheng Yang, Tianci Zhou, Lifan Yuan, Penghao Zhu, Eli Chertkov, Shengyan Liu, Yufeng Du, Ziming Ji, Indranil Das, Qingzhi Chen, Junyi Cao, Yufeng Du, Jiabin Yu, Peixue Wu, Jinchen He, Yifan Su, Yikun Jiang, Yujie Zhang, Chang Liu, Ze-Min Huang, Weizhen Jia, Yunkai Wang, Farshid Jafarpour, Yong Zhao, Xinan Chen, Jessie Shelton, Aaron W. Young, John Bartolotta, Wenchao Xu, Yue Sun, Anjun Chu, Victor Colussi, Chris Akers, Nathan Brooks, Wenbo Fu, Jinchao Zhao, Marvin Qi, Anqi Mu, Yubo Yang, Allen Zang, Yang Lyu, Peizhi Mai, Christopher Wilson, Xuefei Guo, Juntai Zhou, Daniel Inafuku, Chi Xue, Luyu Gao, Ze Yang, Ya\"ir Hein, Yonatan Kahn, Kevin Zhou, Di Luo, John Drew Wilson, Jarrod T. Reilly, Dmytro Bandak, Ofir Press, Liang Yang, Xueying Wang, Hao Tong, Nicolas Chia, Eliu Huerta, Hao Peng 5/12/2026

Probing the Critical Point (CritPt) of AI Reasoning: a Frontier Physics Research Benchmark

CritPt benchmark evaluates LLM reasoning on complex open-ended frontier physics research challenges beyond high-school math and coding.

Ax Adam Tauman Kalai, Yael Tauman Kalai, Or Zamir 5/12/2026

Consensus Sampling for Safer Generative AI

Consensus sampling algorithm aggregates multiple probability distributions to improve generative AI safety with architecture-agnostic approach.

Ax Alex L. Zhang, Tim Kraska, Omar Khattab 5/12/2026

Recursive Language Models

Recursive Language Models enable LLMs to process arbitrarily long prompts through inference-time scaling via recursive self-calling over prompt snippets.

Ax Xuan Yang, Furong Jia, Roy Xie, Xiong Xi, Hengwei Bian, Jian Li, Monica Agrawal 5/12/2026

Batch-of-Thought: Cross-Instance Learning for Enhanced LLM Reasoning

Batch-of-Thought method processes related queries jointly to improve LLM reasoning by identifying high-quality reasoning templates and detecting errors through consistency analysis.

Ax Xingyuan Hua, Sheng Yue, Xinyi Li, Yizhe Zhao, Jinrui Zhang, Ju Ren 5/12/2026

Context Learning for Multi-Agent Discussion

M2CL framework improves multi-agent discussion by learning context representations that align individual LLM instances toward coherent solutions.

Ax Varun Ursekar, Apaar Shanker, Veronica Chatrath, Yuan Xue, Sam Denton 5/12/2026

VeRO: An Evaluation Harness for Agents to Optimize Agents

VeRO evaluation harness systematically measures coding agent performance on agent optimization through iterative edit-execute-evaluate cycles.

Ax Elron Bandel, Asaf Yehudai, Lilach Eden, Yehoshua Sagron, Yotam Perlitz, Elad Venezian, Natalia Razinkov, Natan Ergas, Shlomit Shachor Ifergan, Segev Shlomov, Michal Jacovi, Leshem Choshen, Liat Ein-Dor, Yoav Katz, Michal Shmueli-Scheuer 5/12/2026

General Agent Evaluation

First systematic comparison of agent architectures (tool-calling, MCP, code-generation, CLI) across heterogeneous environments and benchmarks.