Ax Tao Wang, Suhang Zheng, Xiaoxiao Xu 4/14/2026

RTMC: Step-Level Credit Assignment via Rollout Trees

Rollout tree-based credit assignment method for multi-step agentic RL, leveraging implicit state overlap between group rollouts to avoid uniform advantage assignment.

Ax Wei Li, Hangjie Yuan, Zixiang Zhao, Borui Kang, Ziwei Liu, Tao Feng 4/14/2026

A Faster Path to Continual Learning

Optimization technique for continual learning reducing computational overhead of C-Flat while maintaining ability to balance new and old task performance.

Ax Siyu Sun, Jing Ren, Zhaohe Liao, Dongxiao Mao, Xiangyuan Ren, Yiyi Zhang, Haohua Zhao, Weixiong Lin, Jiang Shaohua, Liqing Zhang, Yuchao Zheng 4/14/2026

Bottleneck Tokens for Unified Multimodal Retrieval

Bottleneck tokens framework for unified multimodal retrieval in decoder-only MLLMs, providing explicit pooling and token-level guidance for embedding alignment.

Ax Vikrant Malik, Taylan Kargin, Babak Hassibi 4/14/2026

Distributionally Robust K-Means Clustering

Distributionally robust variant of k-means clustering using Wasserstein-2 balls to protect against outliers, distribution shifts, and limited sample sizes.

Ax Chenhao Fang, Jordi Mola, Mark Harman, Jason Nawrocki, Vaibhav Shrivastava, Yue Cheng, Jay Minesh Shah, Katayoun Zand, Mansi Tripathi, Arya Pudota, Matthew Becker, Herv\'e Robert, Abhishek Gulati 4/14/2026

Reducing Hallucination in Enterprise AI Workflows via Hybrid Utility Minimum Bayes Risk (HUMBR)

Meta's approach to reducing LLM hallucination in enterprise workflows by framing mitigation as Minimum Bayes Risk problem, critical for legal and compliance applications.

Ax Yang Liu, Enxi Wang, Yufei Gao, Weixin Zhang, Bo Wang, Zhiyuan Zeng, Yikai Zhang, Yining Zheng, Xipeng Qiu 4/14/2026

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping

Memory-Enhanced Dynamic reward Shaping (MEDS) framework for reinforcement learning that reduces failure pattern recurrence in LLM training.

Ax Shanchuan Lin, Ceyuan Yang, Zhijie Lin, Hao Chen, Haoqi Fan 4/14/2026

Continuous Adversarial Flow Models

Continuous-time flow models trained with adversarial objectives using learned discriminators instead of fixed MSE criteria.