Ax Jun Liu, Zhenglun Kong, Peiyan Dong, Changdi Yang, Tianqi Li, Hao Tang, Geng Yuan, Wei Niu, Wenbin Zhang, Pu Zhao, Xue Lin, Dong Huang, Yanzhi Wang 3/13/2026

Structured Agent Distillation for Large Language Model

Structured Agent Distillation compresses LLM-based agents into smaller student models while preserving reasoning and action consistency.

Ax Chengyu Shen, Zhen Hao Wong, Runming He, Hao Liang, Meiyi Qiang, Zimo Meng, Zhengyang Zhao, Bohan Zeng, Zhengzhou Zhu, Bin Cui, Wentao Zhang 3/13/2026

Let's Verify Math Questions Step by Step

Framework for verifying correctness of math questions used in LLM training. Focuses on QA data quality beyond answer correctness.

Ax Kai Li, Can Shen, Yile Liu, Jirui Han, Kelong Zheng, Xuechao Zou, Lionel Z. Wang, Shun Zhang, Xingjian Du, Hanjun Luo, Yingbin Jin, Xinxin Xing, Ziyang Ma, Yue Liu, Yifan Zhang, Junfeng Fang, Kun Wang, Yibo Yan, Gelei Deng, Haoyang Li, Yiming Li, Xiaobin Zhuang, Tianlong Chen, Qingsong Wen, Tianwei Zhang, Yang Liu, Haibo Hu, Zhizheng Wu, Xiaolin Hu, Eng-Siong Chng, Wenyuan Xu, XiaoFeng Wang, Wei Dong, Xinfeng Li 3/13/2026

AudioTrust: Benchmarking the Multifaceted Trustworthiness of Audio Large Language Models

AudioTrust benchmark evaluating trustworthiness of audio LLMs. Reveals vulnerabilities from non-semantic acoustic cues like timbre and accent.

Ax Sirui Lu, Zhijing Jin, Terry Jingchen Zhang, Pavel Kos, J. Ignacio Cirac, Bernhard Sch\"olkopf 3/13/2026

Can Theoretical Physics Research Benefit from Language Agents?

Investigates LLM limitations in theoretical physics. Identifies gaps in physical intuition and constraint satisfaction beyond prompting improvements.

Ax Nadav Kunievsky, James A. Evans 3/13/2026

Measuring Intent Comprehension in LLMs

Study measuring how well LLMs comprehend user intent beyond surface-level text matching. Analyzes gap between token prediction and actual user goals.

Ax Zhejun Zhao, Yuchen Li, Alley Liu, Yuehu Dong, Xiaolong Wei, Lixue Zheng, Pingsheng Liu, Dongdong Shen, Long Xia, Jiashu Zhao, Dawei Yin 3/13/2026

TURA: Tool-Augmented Unified Retrieval Agent for AI Search

TURA proposes a tool-augmented retrieval agent for conversational AI search that handles real-time data and structured queries beyond traditional RAG limitations.

Ax Yicheng Di 3/13/2026

LLM-driven Multimodal Recommendation

arXiv: LMMRec framework using LLMs to model user motivations in multimodal recommendation systems from heterogeneous information sources.

Ax Xinwu Ye, Yicheng Mao, Jia Zhang, Yimeng Liu, Li Hao, Fang Wu, Zhiwei Li, Yuxuan Liao, Zehong Wang, Zhiyuan Liu, Zhenfei Yin, Li Yuan, Philip Torr, Huan Sun, Xiangxiang Zeng, Mengdi Wang, Le Cong, Shenghua Gao, Xiangru Tang 3/13/2026

LatentChem: From Textual CoT to Latent Thinking in Chemical Reasoning

arXiv: LatentChem decouples chemical reasoning from natural language by using latent representations instead of chain-of-thought prompting.

Ax Jose Javier Gonzalez Ortiz, Abhay Gupta, Christopher Rinard, Davis Blalock 3/13/2026

FlashOptim: Optimizers for Memory-Efficient Training

arXiv: FlashOptim memory-efficient optimizers for mixed-precision neural network training reducing parameter storage overhead.

Ax Qian Da, Yijiang Chen, Min Ju, Zheyi Ji, Albert Zhou, Wenwen Wang, Matthew A Abikenari, Philip Chikontwe, Guillaume Larghero, Bowen Chen, Peter Neidlinger, Dingrong Zhong, Shuhao Wang, Wei Xu, Drew Williamson, German Corredor, Sen Yang, Le Lu, Xiao Han, Kun-Hsing Yu, Jun-zhou Huang, Laura Barisoni, Geert Litjens, Anant Madabhushi, Lifeng Zhu, Chaofu Wang, Junhan Zhao, Weiguo Hu 3/13/2026

Computational Pathology in the Era of Emerging Foundation and Agentic AI -- International Expert Perspectives on Clinical Integration and Translational Readiness

Expert perspectives on integrating foundation models and AI agents into clinical computational pathology with translational readiness assessment.

Ax Maxwell Miller-Golub, Collin Coil, Kamil Faber, Marcin Pietron, Panpan Zheng, Pasquale Minervini, Roberto Corizzo 3/13/2026

Rethinking the Harmonic Loss via Non-Euclidean Distance Layers

Proposes harmonic loss as alternative to cross-entropy for training deep neural networks with improved interpretability.

Ax Artem Dvirniak, Evgeny Kushnir, Dmitrii Tarasov, Artem Iudin, Oleg Kiriukhin, Mikhail Pautov, Dmitrii Korzh, Oleg Y. Rogov 3/13/2026

Towards Robust Speech Deepfake Detection via Human-Inspired Reasoning

Human-inspired reasoning approach for robust speech deepfake detection with improved generalization to unseen audio domains and interpretability.