Ax Zhaoyu Zhu, Shuhan Zhang, Rui Gao, Shuang Li 3/4/2026

Wasserstein Proximal Policy Gradient

Derives Wasserstein Proximal Policy Gradient using optimal transport geometry for continuous-action entropy-regularized RL without policy log-density evaluation.

Ax Yunxiang Li, Mark Schmidt, Reza Babanezhad, Sharan Vaswani 3/4/2026

Towards Parameter-Free Temporal Difference Learning

Develops parameter-free temporal difference learning for RL that avoids requiring problem-dependent quantities like feature covariance eigenvalues.

Ax Zhixia Zhang, Zixuan Huang, Xin Xia, Deqing Wang, Fuzhen Zhuang, Shuai Ma, Ning Ding, Yaodong Yang, Jianxin Li, Yikun Ban 3/4/2026

Heterogeneous Agent Collaborative Reinforcement Learning

Introduces HACRL, a collaborative reinforcement learning paradigm where heterogeneous agents share verified rollouts during training but execute independently at inference.

Ax Ryan Feng Lin, Yuantao Wei, Huiling Liao, Xiaoning Qian, Shuai Huang 3/4/2026

Causal Learning Should Embrace the Wisdom of the Crowd

Paradigm for causal structure learning from observational data leveraging human causal knowledge to address combinatorial explosion of possible graphs.

Ax George Bredis, Nikita Balagansky, Daniil Gavrilov, Ruslan Rakhimov 3/4/2026

Next Embedding Prediction Makes World Models Stronger

NE-Dreamer agent uses temporal transformer to predict next-step embeddings for improved model-based reinforcement learning in high-dimensional domains.

Ax Zhenquan Yao, Zitong Huang, Yihan Zeng, Jianhua Han, Hang Xu, Chun-Mei Feng, Jianwei Ma, Wangmeng Zuo 3/4/2026

CGL: Advancing Continual GUI Learning via Reinforcement Fine-Tuning

Framework for continual learning in GUI agents using multimodal LLMs with reinforcement fine-tuning to adapt to new tasks without catastrophic forgetting.

Ax Robin Young 3/4/2026

Why Does RLAIF Work At All?

arXiv: Theoretical explanation for reinforcement learning from AI feedback through latent value hypothesis.

Ax Dan Stowell 3/4/2026

Torus embeddings

Torus embeddings: research on representing deep learning embeddings on toroidal manifolds instead of Euclidean space for efficiency.