Ax Mehdi Acheli, Walid Gaaloul 2/25/2026

Motivation is Something You Need

Proposes dual-model training framework inspired by neuroscience motivation states with alternating base and larger model activation.

Ax Debjit Paul, Daniel Murphy, Milan Gritta, Ronald Cardenas, Victor Prokhorov, Lena Sophia Bolliger, Aysim Toker, Roy Miles, Andreea-Maria Oncescu, Jasivan Alex Sivakumar, Philipp Borchert, Ismail Elezi, Meiru Zhang, Ka Yiu Lee, Guchun Zhang, Jun Wang, Gerasimos Lampouras 2/25/2026

A Benchmark for Deep Information Synthesis

DEEPSYNTH benchmark evaluates LLM-based agents on complex multi-source information synthesis tasks beyond fact retrieval.

Ax Tony Feng, Junehyuk Jung, Sang-hyun Kim, Carlo Pagano, Sergei Gukov, Chiang-Chiang Tsai, David Woodruff, Adel Javanmard, Aryan Mokhtari, Dawsen Hwang, Yuri Chervonyi, Jonathan N. Lee, Garrett Bingham, Trieu H. Trinh, Vahab Mirrokni, Quoc V. Le, Thang Luong 2/25/2026

Aletheia tackles FirstProof autonomously

Aletheia, a mathematics research AI agent powered by Gemini 3 Deep Think, autonomously solves 6 of 10 FirstProof challenge problems.

Ax Yebo Wu, Chunlin Tian, Jingguang Li, He Sun, Kahou Tam, Zhanting Zhou, Haicheng Liao, Jing Xiong, Zhijiang Guo, Li Li, Chengzhong Xu 2/25/2026

A Survey on Federated Fine-tuning of Large Language Models

Comprehensive survey of Federated Learning combined with LLM fine-tuning (FedLLM), covering privacy-preserving collaborative model adaptation methods.

Ax Yucheng Shi, Wenhao Yu, Jingyuan Huang, Wenlin Yao, Wenhu Chen, Ninghao Liu 2/25/2026

Towards Trustworthy GUI Agents: A Survey

Survey of trustworthy GUI agents built on LLMs, identifying execution gap challenges in real-world digital environment automation with irreversible actions.

Ax Jing Yu Lim, Rushi Shah, Zarif Ikram, Samson Yu, Haozhe Ma, Tze-Yun Leong, Dianbo Liu 2/25/2026

Performance Asymmetry in Model-Based Reinforcement Learning

Analysis of performance asymmetry in Model-Based RL agents on Atari100k, showing dramatic variance across task types despite high average performance.

Ax Zahra Shahrooei, Ali Baheri 2/25/2026

Wasserstein Barycenter Soft Actor-Critic

Wasserstein Barycenter Soft Actor-Critic algorithm improves sample efficiency in off-policy reinforcement learning via directed exploration.

Ax Andrey Goncharov, Daniil Vyazhev, Petr Sychev, Edvard Khalafyan, Alexey Zaytsev 2/25/2026

Complexity-aware fine-tuning

Efficient fine-tuning method for LLMs using entropy-based complexity detection to apply chain-of-thought reasoning selectively on difficult examples.

Ax Xuefeng Liu, Mingxuan Cao, Songhao Jiang, Xiao Luo, Xiaotian Duan, Mengdi Wang, Tobin R. Sosnick, Jinbo Xu, Rick Stevens 2/25/2026

Monte Carlo Tree Diffusion with Multiple Experts for Protein Design

MCTD-ME combines masked diffusion models with Monte Carlo Tree Search for protein design, addressing long-range dependencies and search space challenges.

Ax Jubayer Ibn Hamid, Ifdita Hasan Orney, Ellen Xu, Chelsea Finn, Dorsa Sadigh 2/25/2026

Polychromic Objectives for Reinforcement Learning

Polychromic objectives framework for reinforcement learning fine-tuning preserves policy diversity during RLFT to prevent mode collapse.

Ax Siddarth Venkatraman, Vineet Jain, Sarthak Mittal, Vedant Shah, Johan Obando-Ceron, Yoshua Bengio, Brian R. Bartoldson, Bhavya Kailkhura, Guillaume Lajoie, Glen Berseth, Nikolay Malkin, Moksh Jain 2/25/2026

Recursive Self-Aggregation Unlocks Deep Thinking in Large Language Models

Recursive Self-Aggregation (RSA) test-time scaling method combines parallel and sequential inference to improve LLM reasoning capabilities.

Ax Lizhang Chen, Jonathan Li, Kaizhao Liang, Baiyu Su, Cong Xie, Nuo Wang Pierse, Chen Liang, Ni Lao, Qiang Liu 2/25/2026

Cautious Weight Decay

Cautious Weight Decay (CWD) optimizer modification applies weight decay only to parameters aligned with optimizer updates.

Ax Dario Shariatian, Alain Durmus, Umut Simsekli, Stefano Peluchetti 2/25/2026

Latent-Augmented Discrete Diffusion Models

Latent-Augmented Discrete Diffusion (LADD) improves discrete diffusion models for fast language generation by modeling cross-token dependencies.

Ax Alexandra Volkova, Mher Safaryan, Christoph H. Lampert, Dan Alistarh 2/25/2026

Towards Robust Scaling Laws for Optimizers

Research on scaling laws for LLM pretraining with different optimizers beyond AdamW, examining new optimizers like Muon, Shampoo, and SOAP.

Ax Cl\'audio Correia, Alberto E. A. Ferreira, Lucas Martins, Miguel P. Bento, Sofia Guerreiro, Ricardo Ribeiro Pereira, Ana Sofia Gomes, Jacopo Bono, Hugo Ferreira, Pedro Bizarro 2/25/2026

MUSE: Multi-Tenant Model Serving With Seamless Model Updates

Multi-tenant ML serving system handling seamless model updates while maintaining decision thresholds across clients with distribution shifts.

Ax DatologyAI, :, Aldo Gael Carranza, Kaleigh Mentzer, Ricardo Pio Monti, Alex Fang, Alvin Deng, Amro Abbas, Anshuman Suri, Brett Larsen, Cody Blakeney, Darren Teh, David Schwab, Diego Kiner, Fan Pan, Haakon Mongstad, Haoli Yin, Jack Urbanek, Jason Lee, Jason Telanoff, Josh Wills, Luke Merrick, Maximilian B\"other, Parth Doshi, Paul Burstein, Pratyush Maini, Rishabh Adaiga, Sid Joshi, Spandan Das, Tony Jiang, Vineeth Dorna, Zhengping Wang, Bogdan Gaza, Ari Morcos, Matthew Leavitt 2/25/2026

\"UberWeb: Insights from Multilingual Curation for a 20-Trillion-Token Dataset

Study of multilingual data curation across 13 languages identifying interference patterns and optimal training strategies for 20-trillion-token dataset.