Ax Sarvesh Patil, Mitsuhiko Nakamoto, Manan Agarwal, Shashwat Saxena, Jesse Zhang, Giri Anantharaman, Cleah Winston, Chaoyi Pan, Douglas Chen, Nai-Chieh Huang, Zeynep Temel, Oliver Kroemer, Sergey Levine, Abhishek Gupta, Hongkai Da, Paarth Shah, Max Simchowitz 5/6/2026

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies

Sample-efficient fine-tuning algorithm for diffusion and flow-based robot control policies using off-policy critic networks and modified PPO.

Ax Prakhar Gupta, Garv Shah, Donghua Zhang 5/6/2026

Self-Mined Hardness for Safety Fine-Tuning

Self-Mined Hardness method for LLM safety fine-tuning that scores prompt difficulty by model jailbreak frequency, then trains on hardest prompts.

Ax Anna Sokol, Marianna B. Ganapini, Nitesh V. Chawla 5/6/2026

Do LLMs have core beliefs?

Examines whether LLMs exhibit core beliefs—foundational commitments that resist change—parallel to human cognition.

Ax Doojin Baek, Gyubin Lee, Junyeob Baek, Hosung Lee, Sungjin Ahn 5/6/2026

Learning to Theorize the World from Observation

Proposes learning-to-theorize approach where agents build internal theories from observations inspired by cognitive science.

Ax Wenjin Hou, Shangpin Peng, Weinong Wang, Zheng Ruan, Yue Zhang, Zhenglin Zhou, Mingqi Gao, Yifei Chen, Kaiqi Wang, Hongming Yang, Chengquan Zhang, Zhuotao Tian, Han Hu, Yi Yang, Fei Wu, Hehe Fan 5/6/2026

Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe

Uni-OPD framework unifies on-policy distillation theory, identifying bottlenecks in consolidating expert models into single student models.