Ax Licheng Pan, Haochen Yang, Haoxuan Li, Yunsheng Lu, Yongqi Tong, Yinuo Wang, Shijian Wang, Zhixuan Chu, Lei Shen, Yuan Lu, Hao Wang 5/8/2026

Optimal Transport for LLM Reward Modeling from Noisy Preference

SelectiveRM framework using optimal transport to handle noisy preferences in LLM reward model training for RLHF.

Ax Maxim Fishman, Brian Chmiel, Ron Banner, Daniel Soudry, Boris Ginsburg 5/8/2026

Normalized Architectures are Natively 4-Bit

nGPT architecture with normalized weights and activations on unit hypersphere enables stable 4-bit precision training without random transforms or scaling tricks.

Ax Zixuan Wang, Yuchen Yan, Hongxing Li, Teng Pan, Dingming Li, Ruiqing Zhang, Weiming Lu, Jun Xiao, Yueting Zhuang, Yongliang Shen 5/8/2026

Milestone-Guided Policy Learning for Long-Horizon Language Agents

Milestone-Guided Policy Learning for long-horizon language agents addresses credit misattribution and sample inefficiency through intermediate milestone supervision.

Ax Hao Lin, Kunyang Lv, Xu Jiang, Jingqi Tian, Zhongjing Du, Jiayu Ding, Qiaoman Zhang, Hongbo Jin 5/8/2026

VISD: Enhancing Video Reasoning via Structured Self-Distillation

VISD enhances VideoLLMs for complex reasoning combining RL with verifiable rewards and structured self-distillation for fine-grained credit assignment.

Ax Bowen Zheng, Weijian Luo, Guang Yang, Colin Zhang, Tianyang Hu 5/8/2026

Autoregressive Visual Generation Needs a Prologue

Prologue approach for autoregressive image generation prepending learnable prologue tokens to bridge reconstruction-generation gap in visual token sequences.

Ax Samir Darouich, Vinh Tong, Llu\'is Pastor-P\'erez, Tanja Bien, Loay Mualem, Mathias Niepert 5/8/2026

SymDrift: One-Shot Generative Modeling under Symmetries

SymDrift approach for one-shot generative modeling of physical systems using equivariant diffusion models that respect global symmetries like rotations.

Ax Ajay Jaiswal, Lauren Hannah, Han-Byul Kim, Duc Hoang, Mehrdad Farajtabar, Minsik Cho 5/8/2026

TIDE: Every Layer Knows the Token Beneath the Context

Research on improving LLM architecture by using token indices at every layer instead of once at input, addressing rare token training and position awareness issues.