Ax Vedant Shah, Johan Obando-Ceron, Vineet Jain, Brian Bartoldson, Bhavya Kailkhura, Sarthak Mittal, Glen Berseth, Pablo Samuel Castro, Yoshua Bengio, Nikolay Malkin, Moksh Jain, Siddarth Venkatraman, Aaron Courville 3/19/2026

A Comedy of Estimators: On KL Regularization in RL Training of LLMs

Analysis of KL divergence estimators used in RL training of LLMs, evaluating different approximation methods for reverse KL regularization.

Ax Abhi Kottamasu, Chirag Mahapatra, Sam Lee, Ben Pan, Aakash Barthwal, Akul Datta, Ajay Arun, Silas Alberti, Adarsh Hiremath, Brendan Foody, Bertie Vidgen 3/19/2026

APEX-SWE

Benchmark for evaluating frontier AI models on real-world software engineering tasks including integration and end-to-end system construction.

Ax Xiaoxuan Liu, Jiaxiang Yu, Jongseok Park, Ion Stoica, Alvin Cheung 3/19/2026

Speculative Decoding: Performance or Illusion?

Systematic evaluation of speculative decoding acceleration techniques on production-grade vLLM inference engine, assessing real-world effectiveness.

Ax Yuanhe Zhang, Xinyue Wang, Zhican Chen, Weiliu Wang, Zilu Zhang, Zhengshuo Gong, Zhenhong Zhou, Kun Wang, Li Sun, Yang Liu, Sen Su 3/19/2026

Resource Consumption Threats in Large Language Models

Survey of resource consumption threats in LLMs, covering efficiency issues affecting service capacity, latency, and API costs.

Ax Zhengbo Zhang, Jinbo Su, Zhaowen Zhou, Changtao Miao, Yuhan Hong, Qimeng Wu, Yumeng Liu, Feier Wu, Yihe Tian, Yuhao Liang, Zitong Shan, Wanke Xia, Yi-Fan Zhang, Bo Zhang, Zhe Li, Shiming Xiang, Ying Yan 3/19/2026

VisBrowse-Bench: Benchmarking Visual-Native Search for Multimodal Browsing Agents

Benchmark for visual-native search in multimodal browsing agents, evaluating MLLM visual reasoning over web pages.

Ax G. Ciarfaglia, A. Rosanova, S. Cipolla, J. Bartoli, A. Di Domenico, C. Fioroni, A. Fontana, M. R. Scoleri, M. I. Mone, D. Franchi, M. C. Del Gaudio, F. Picariello, M. Gabusi, S. Bonura, V. Morreale, I. Bailo 3/19/2026

EngGPT2: Sovereign, Efficient and Open Intelligence

Italian open-source LLM with 16B parameters achieving competitive performance on benchmarks while requiring fraction of inference power.

Ax Omer Nacar, Deema Alquffari, Saleh Alsharideh, Adeem AlOtaibi, Abdulaziz Alabdulkarim, Leen Alhazmi, Nada Alomar, Wareef Alzubaidi, Nada Alsultan, Ahmed Alrabghi, Demah Alhoshan, Rana Alsayyari, Hamed Alruwaili, Albaraa Jaafar, Khaled Alusmani, Abdulaziz Alsohimy, Munirah Alsubaie, Shahd Aldukhayil, Arwa Alali, Yazeed BinShihah, Razan Alsulaymi, Nourah Alhumaid, Razan Abdulsalam, Reem Alamoudi, Mohammed Alkhalifa 3/19/2026

From Language to Action in Arabic: Reliable Structured Tool Calling via Data-Centric Fine-Tuning

Production framework for reliable Arabic function-calling models enabling agentic AI systems through data-centric fine-tuning.

Ax Benjamin Hudson, Laurent Charlin, Emma Frejinger 3/19/2026

Contextual Preference Distribution Learning

Proposes pipeline to learn context-dependent preference distributions for risk-averse decision-making via inverse optimization.