Ax Guojiang Zhao, Zixiang Lu, Yutang Ge, Sihang Li, Zheng Cheng, Haitao Lin, Lirong Wu, Hanchen Xia, Hengxing Cai, Wentao Guo, Hongshuai Wang, Mingjun Xu, Siyu Zhu, Guolin Ke, Linfeng Zhang, Zhifeng Gao 2/24/2026

MolReasoner: Toward Effective and Interpretable Reasoning for Molecular LLMs

MolReasoner framework for domain-specific molecular reasoning in LLMs using specialized prompting to reduce hallucinations.

Ax Jubayer Ibn Hamid, Ifdita Hasan Orney, Ellen Xu, Chelsea Finn, Dorsa Sadigh 2/24/2026

Polychromic Objectives for Reinforcement Learning

Polychromic objectives method for reinforcement learning fine-tuning preventing policy collapse and preserving diversity during RLFT training.

Ax Santiago Cuervo, Skyler Seto, Maureen de Seyssel, Richard He Bai, Zijin Gu, Tatiana Likhomanenko, Navdeep Jaitly, Zakaria Aldeneh 2/24/2026

Closing the Gap Between Text and Speech Understanding in LLMs

Analyzes performance gap between speech-adapted and text-based LLMs on language understanding tasks, identifying causes of speech input underperformance.

Ax Seohong Park, Aditya Oberai, Pranav Atreya, Sergey Levine 2/24/2026

Transitive RL: Value Learning via Divide and Conquer

Transitive RL: Divide-and-conquer value learning algorithm for offline goal-conditioned reinforcement learning using triangle inequality structure.

Ax Xiao Wu, Ting-Zhu Huang, Liang-Jian Deng, Xiaobing Yu, Yu Zhong, Shangqi Deng, Ufaq Khan, Jianghao Wu, Xiaofeng Liu, Imran Razzak, Xiaojun Chang, Yutong Xie 2/24/2026

SelfAI: A self-directed framework for long-horizon scientific discovery

SelfAI: Multi-agent system for long-horizon scientific discovery with self-directed exploration, balancing efficiency-diversity trade-offs in complex hypothesis spaces.

Ax Yaswanth Chittepu, Raghavendra Addanki, Tung Mai, Anup Rao, Branislav Kveton 2/24/2026

ML-Tool-Bench: Tool-Augmented Planning for ML Tasks

ML-Tool-Bench: Framework for autonomous ML agents using LLMs to orchestrate end-to-end data science workflows including analysis, feature engineering, and hyperparameter optimization.

Ax Md. Najib Hasan (Wichita State University, USA), Imran Ahmad (Wichita State University, USA), Sourav Basak Shuvo (Khulna University of Engineering and Technology, Bangladesh), Md. Mahadi Hasan Ankon (Khulna University of Engineering and Technology, Bangladesh), Sunanda Das (University of Arkansas, USA), Nazmul Siddique (Ulster University, UK), Hui Wang (Queen's University Belfast, UK) 2/24/2026

DL$^3$M: A Vision-to-Language Framework for Expert-Level Medical Reasoning through Deep Learning and Large Language Models

Vision-to-language framework combining medical image classification with LLM reasoning for interpretable clinical decision support.