Ax Seth Karten, Jake Grigsby, Tersoo Upaa Jr, Junik Bae, Seonghun Hong, Hyunyoung Jeong, Jaeyoon Jung, Kun Kerdthaisong, Gyungbo Kim, Hyeokgi Kim, Yujin Kim, Eunju Kwon, Dongyu Liu, Patrick Mariglia, Sangyeon Park, Benedikt Schink, Xianwei Shi, Anthony Sistilli, Joseph Twin, Arian Urdu, Matin Urdu, Qiao Wang, Ling Wu, Wenli Zhang, Kunsheng Zhou, Stephanie Milani, Kiran Vodrahalli, Amy Zhang, Fei Fang, Yuke Zhu, Chi Jin 3/17/2026

The PokeAgent Challenge: Competitive and Long-Context Learning at Scale

PokeAgent Challenge: Large-scale benchmark for competitive multi-agent decision-making with partial observability and long-horizon planning.

Ax Lianghui Zhu, Yuxin Fang, Bencheng Liao, Shijie Wang, Tianheng Cheng, Zilong Huang, Chen Chen, Lai Wei, Yutao Zeng, Ya Wang, Yi Lin, Yu Li, Xinggang Wang 3/17/2026

Mixture-of-Depths Attention

Mixture-of-Depths Attention: Mechanism addressing signal degradation in deep LLMs by enabling attention to multiple depth levels.

Ax Sydney Levine, Matija Franklin, Tan Zhi-Xuan, Secil Yanik Guyot, Lionel Wong, Daniel Kilov, Yejin Choi, Joshua B. Tenenbaum, Noah Goodman, Seth Lazar, Iason Gabriel 3/17/2026

Resource Rational Contractualism Should Guide AI Alignment

Framework for AI alignment grounded in resource-rational contractualism, enabling diverse stakeholders to reach agreements on AI decision-making.

Ax Monoshiz Mahbub Khan, Xiaoyin Xi, Andrew Meneely, Yiming Tang, Zhe Yu 3/17/2026

Efficient Story Point Estimation With Comparative Learning

Machine learning approach to automate story point estimation for software sprint planning using comparative learning from historical team decisions.

Ax Zhehao Dong, Xiaofeng Wang, Zheng Zhu, Yirui Wang, Yang Wang, Yukun Zhou, Boyuan Wang, Chaojun Ni, Runqi Ouyang, Wenkang Qin, Xinze Chen, Yun Ye, Guan Huang, Zhen Lu, Yue Yang 3/17/2026

EMMA: Generalizing Real-World Robot Manipulation via Generative Visual Transfer

Data augmentation framework for vision-language-action models in robot manipulation using generative visual transfer to reduce annotation costs.

Ax Maximilian N\"agele, Florian Marquardt 3/17/2026

Agentic Exploration of Physics Models

Agentic framework for automated scientific discovery that iteratively explores unknown systems through experiments and analysis without domain-specific tailoring.

Ax Yichao Liang, Dat Nguyen, Cambridge Yang, Tianyang Li, Joshua B. Tenenbaum, Carl Edward Rasmussen, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis 3/17/2026

ExoPredicator: Learning Abstract Models of Dynamic Worlds for Robot Planning

Framework for learning abstract world models that jointly represent symbolic states and causal processes for endogenous and exogenous dynamics in robot planning.

Ax H M Quamran Hasan, Housam Khalifa Bashier, Jiayi Dai, Mi-Young Kim, Randy Goebel 3/17/2026

Reason2Decide: Rationale-Driven Multi-Task Learning

Reason2Decide: two-stage training framework for clinical decision support LLMs to generate predictions with self-aligned explanations.

Ax Minhua Lin, Hanqing Lu, Zhan Shi, Bing He, Rui Mao, Zhiwei Zhang, Zongyu Wu, Xianfeng Tang, Hui Liu, Zhenwei Dai, Xiang Zhang, Suhang Wang, Benoit Dumoulin, Jian Pei 3/17/2026

Position: Agentic Evolution is the Path to Evolving LLMs

Position paper arguing agentic evolution via deployment-time adaptation is needed to close the train-deploy gap in LLM systems.

Ax Tony Feng, Junehyuk Jung, Sang-hyun Kim, Carlo Pagano, Sergei Gukov, Chiang-Chiang Tsai, David Woodruff, Adel Javanmard, Aryan Mokhtari, Dawsen Hwang, Yuri Chervonyi, Jonathan N. Lee, Garrett Bingham, Trieu H. Trinh, Vahab Mirrokni, Quoc V. Le, Thang Luong 3/17/2026

Aletheia tackles FirstProof autonomously

Aletheia mathematics research agent solved 6 of 10 FirstProof challenge problems autonomously using Gemini 3 Deep Think reasoning.

Ax Chen Bo Calvin Zhang, Christina Q. Knight, Nicholas Kruus, Jason Hausenloy, Pedro Medeiros, Nathaniel Li, Aiden Kim, Yury Orlovskiy, Coleman Breen, Bryce Cai, Jasper G\"otting, Andrew Bo Liu, Samira Nedungadi, Paula Rodriguez, Yannis Yiming He, Mohamed Shaaban, Zifan Wang, Seth Donoughe, Julian Michael 3/17/2026

LLM Novice Uplift on Dual-Use, In Silico Biology Tasks

Human study measuring whether LLM access improves novice performance on biology tasks versus internet-only baselines, with dual-use risk implications.

Ax Shiya Zhang, Yuhan Zhan, Ruixi Su, Ruihan Sun, Ziyi Song, Zhaohan Chen, Xiaofan Zhang 3/17/2026

EMPA: Evaluating Persona-Aligned Empathy as a Process

EMPA framework evaluates how well LLM dialogue agents maintain persona-aligned empathy across multi-turn conversations using process-oriented metrics.

Ax Ann Yuan, Asma Ghandeharioun, Carter Blum, Alicia Machado, Jessica Hoffmann, Daphne Ippolito, Martin Wattenberg, Lucas Dixon, Katja Filippova 3/17/2026

Think Before You Lie: How Reasoning Leads to Honesty

Study showing reasoning and deliberation increase honesty in LLM responses on moral trade-off scenarios.