Ax Peng Xia, Jianwen Chen, Xinyu Yang, Haoqin Tu, Jiaqi Liu, Kaiwen Xiong, Siwei Han, Shi Qiu, Haonian Ji, Yuyin Zhou, Zeyu Zheng, Cihang Xie, Huaxiu Yao 3/19/2026

MetaClaw: Just Talk -- An Agent That Meta-Learns and Evolves in the Wild

MetaClaw enables LLM agents to continuously adapt and evolve in production by meta-learning from diverse task distributions without storing raw trajectories.

Ax Seyed Mohammad Asghari, Chris Chute, Vikranth Dwaracherla, Xiuyuan Lu, Mehdi Jafarnia, Victor Minden, Zheng Wen, Benjamin Van Roy 3/19/2026

Efficient Exploration at Scale

arXiv: Online learning algorithm for RLHF that improves data efficiency. Incrementally updates reward and language models from choice data.

Ax Dilxat Muhtar, Jiashun Liu, Wei Gao, Weixun Wang, Shaopan Xiong, Ju Huang, Siran Yang, Wenbo Su, Jiamang Wang, Ling Pan, Bo Zheng 3/19/2026

Complementary Reinforcement Learning

arXiv paper on complementary reinforcement learning for LLM-based agents, improving sample efficiency by leveraging historical experience across episodes.

Ax Ting Gao, Stavros Orfanoudakis, Nan Lin, Elvin Isufi, Winnie Daamen, Serge Hoogendoorn 3/19/2026

Flow Matching Policy with Entropy Regularization

arXiv paper on flow matching policies with entropy regularization for diffusion-based reinforcement learning, improving policy gradient computation.

Ax Yihong Chen, Quanming Yao 3/19/2026

Attention Sinks Induce Gradient Sinks

arXiv paper studying attention sinks in Transformers from backpropagation perspective, showing attention sinks induce gradient concentration under causal masking.

Ax Luca Hinkamp, Simon Kl\"uttermann, Emmanuel M\"uller 3/19/2026

RangeAD: Fast On-Model Anomaly Detection

RangeAD leverages primary model's learned representations for efficient on-model anomaly detection without separate AD model.

Ax Alexander D. Goldie, Zilin Wang, Adrian Hayler, Deepak Nathani, Edan Toledo, Ken Thampiratwong, Aleksandra Kalisz, Michael Beukman, Alistair Letcher, Shashank Reddy, Clarisse Wibault, Theo Wolf, Charles O'Neill, Uljad Berdica, Nicholas Roberts, Saeed Rahmani, Hannah Erlebach, Roberta Raileanu, Shimon Whiteson, Jakob N. Foerster 3/19/2026

Procedural Generation of Algorithm Discovery Tasks in Machine Learning

DiscoGen procedural generator creates diverse algorithm discovery tasks to improve evaluation of AutoML systems and algorithm design optimization.

Ax Cristiano Capone, Luca Falorsi, Andrea Ciardiello, Luca Manneschi 3/19/2026

Unified Policy Value Decomposition for Rapid Adaptation

Framework for rapid adaptation in reinforcement learning where policy and value functions share low-dimensional embeddings for novel task generalization.