Ax Tara Bogavelli, Gabrielle Gauthier Melan\c{c}on, Katrina Stankiewicz, Oluwanifemi Bamgbose, Fanny Riols, Hoang H. Nguyen, Raghav Mehndiratta, Lindsay Devon Brin, Joseph Marinier, Hari Subramani, Anil Madamala, Sridhar Krishna Nemala, Srinivas Sunkara 5/14/2026

EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents

EVA-Bench: Evaluation framework for voice agents addressing realistic conversation simulation and voice-specific failure mode measurement.

Ax Terry Jingchen Zhang, Gopal Dev, Ning Wang, Max Obreiter, Punya Syon Pandey, Keenan Samway, Wenyuan Jiang, Yinya Huang, Bernhard Sch\"olkopf, Mrinmaya Sachan, Zhijing Jin 5/14/2026

Test of Time: Rethinking Temporal Signal of Benchmark Contamination

Critical analysis of temporal signals in benchmark contamination detection, showing sensitivity to question construction independent of data memorization.

Ax Youngmin Im, Byeongung Jo, Jaeyoung Wi, Seungwoo Baek, Tae Hoon Min, Joo Hyung Lee, Sangeun Oh, Insik Shin, Sunjae Lee 5/14/2026

MobiBench: Multi-Branch, Modular Benchmark for Mobile GUI Agents

MobiBench: Multimodal benchmark for mobile GUI agents addressing limitations of existing offline/online benchmarks with multiple valid action paths.

Ax Sitao Cheng, Tianle Li, Xuhan Huang, Xunjian Yin, Difan Zou 5/14/2026

Differentiable Evolutionary Reinforcement Learning

Differentiable Evolutionary RL optimizes reward functions using gradient information to improve policy performance on complex reasoning tasks.

Ax Seanie Lee, Sangwoo Park, Yumin Choi, Gyeongman Kim, Minki Kang, Jihun Yun, Dongmin Park, Jongho Park, Sung Ju Hwang 5/14/2026

THINKSAFE: Self-Generated Safety Alignment for Reasoning Models

THINKSAFE: Safety alignment approach for reasoning models that self-generates safety constraints without external teacher distillation.

Ax Xingyuan Hua, Sheng Yue, Xinyi Li, Yizhe Zhao, Jinrui Zhang, Ju Ren 5/14/2026

Context Learning for Multi-Agent Discussion

M2CL: Multi-LLM context learning method for multi-agent discussion systems addressing discussion inconsistency and context misalignment.

Ax Tae Soo Kim, Yoonjoo Lee, Jaesang Yu, John Joon Young Chung, Juho Kim 5/14/2026

DiscoverLLM: From Executing Intents to Discovering Them

DiscoverLLM: Method enabling LLMs to help users discover intents through interactive exploration rather than just executing stated requests.

Ax Baoqing Yue, Zihan Zhu, Yutong Han, Qian Sun, Jichen Feng, Hufei Yang, Yifan Zhang, Mengdi Wang 5/14/2026

Interactive Benchmarks

Interactive Benchmarks: Evaluation paradigm assessing model reasoning by testing ability to decide what information to acquire and use.

Ax Wanyi Chen, Xiao Yang, Xu Yang, Tianming Sha, Qizheng Li, Zhuo Wang, Bowen Xian, Fang Kong, Weiqing Liu, Jiang Bian 5/14/2026

Agent^2 RL-Bench: Can LLM Agents Engineer Agentic RL Post-Training?

Benchmark evaluating whether LLM agents can autonomously design, implement, and execute RL post-training pipelines for model improvement.

Ax Aleksandr Bowkis, Marie Davidsen Buhl, Jacob Pfau, Geoffrey Irving 5/14/2026

Automated alignment is harder than you think

Paper arguing automated alignment via research agents risks producing misleading safety assessments without deliberate sabotage.

Ax Daniel Zheng, Ingrid von Glehn, Yori Zwols, Iuliya Beloshapka, Lars Buesing, Daniel M. Roy, Martin Wattenberg, Bogdan Georgiev, Tatiana Schmidt, Andrew Cowie, Fernanda Viegas, Dimitri Kanevsky, Vineet Kahlon, Hartmut Maennel, Sophia Alj, George Holland, Alex Davies, Pushmeet Kohli 5/14/2026

AI co-mathematician: Accelerating mathematicians with agentic AI

AI co-mathematician workbench enabling mathematicians to collaboratively leverage AI agents for research including ideation, computation, and theorem proving.

Ax Shawn Li, Chenxiao Yu, Han Wang, Wei Yang, Ryan Rossi, Franck Dernoncourt, Xiyang Hu, Philip Yu, Chaowei Xiao, Huan Zhang, Yue Zhao 5/14/2026

FORTIS: Benchmarking Over-Privilege in Agent Skills

FORTIS benchmark evaluating privilege escalation vulnerabilities in LLM agent skill layers and access control policies.

Ax Songlin Bai, Xintong Wang, Linlin Yu, Bin Chen, Zhiang Xu, Yuyang Sheng, Changtong Zan, Xiaofeng Zhu, Yizhe Zhang, Jiru Li, Mingze Guo, Ling Zou, Yalong Li, Chengfu Huo, Liang Ding 5/14/2026

IndustryBench: Probing the Industrial Knowledge Boundaries of LLMs

IndustryBench: 2,049-item benchmark testing LLM knowledge on industrial procurement with safety-critical constraints and standards compliance.

Ax Peter Baile Chen, Devin Yang, Weiyue Li, Fabian Wenz, Yi Zhang, Nesime Tatbul, Michael Cafarella, \c{C}a\u{g}atay Demiralp, Michael Stonebraker 5/14/2026

BEAVER: An Enterprise Benchmark for Text-to-SQL

Benchmark for evaluating LLM text-to-SQL performance on complex enterprise databases with intricate schemas and domain knowledge requirements.