Ax Aaditya L. Kachhadiya 5/14/2026

Local Inverse Geometry Can Be Amortized

Amortized learned surrogate (Deceptron) for solving nonlinear inverse problems by encoding curvature information into reusable reverse operator.

Ax Huiqi Deng, Yibo Li, Quanshi Zhang, Peng Zhang, Hongbin Pei, Xia Hu 5/14/2026

Understanding Generalization through Decision Pattern Shift

Decision Pattern Shift framework analyzing how deep neural network internal decision mechanisms evolve from training to test data for understanding generalization failures.

Ax Xinyu Liu, Kechen Jiao, Chunyang Xiao, Runsong Zhao, Junhao Ruan, Bei Li, Jiahao Liu, Qifan Wang, Xin Chen, Jingang Wang, Tong Xiao, JingBo Zhu 5/14/2026

Teacher-Guided Policy Optimization for LLM Distillation

arXiv paper: Teacher-Guided Policy Optimization for LLM distillation using Reverse KL to improve student-teacher convergence.

Ax Ian Osband 5/14/2026

Delightful Exploration

Exploration algorithm balancing uncertainty resolution with action budget using expected improvement and surprisal gating.

Ax Jan Arne Telle, Brigt H{\aa}vardstun, Jose Hernandez-Orallo 5/14/2026

Teaching and Learning under Deductive Errors

Teaching framework for machine learning that accounts for deductive errors in learners like LLMs during few-shot learning.

Ax Akhil Premkumar, Sarah Lucioni 5/14/2026

The Diffusion Encoder

Novel encoder architecture replacing VAE encoders with diffusion models for improved latent representation learning.

Ax Neeratyoy Mallik, Maciej Janowski, Johannes Hog, Herilalaina Rakotoarison, Josif Grabocka, Frank Hutter, Aaron Klein 5/14/2026

When is Warmstarting Effective for Scaling Language Models?

Research on warmstarting techniques for scaling language models, analyzing initialization constraints and growth strategies for training efficiency.

Ax Asim Osman, Sasha Abramowitz, Mark Bergh, Ulrich Armel Mbou Sob, Ruan John de Kock, Omayma Mahjoub, Oussama Hidaoui, Noah De Nicola, Arnol Manuel Fokam, Felix Chalumeau, Daniel Rajaonarivonivelomanantsoa, Siddarth Singh, Refiloe Shabe, Juan Claude Formanek, Simon Verster Du Toit, Arnu Pretorius 5/14/2026

Self-Supervised On-Policy Reinforcement Learning via Contrastive Proximal Policy Optimisation

Self-supervised contrastive reinforcement learning algorithm for discrete action spaces without hand-crafted rewards.

Ax Yuetai Li, Fengqing Jiang, Yichen Feng, Kaiyuan Zheng, Luyao Niu, Bhaskar Ramasubramanian, Basel Alomair, Linda Bushnell, Radha Poovendran 5/14/2026

Polyhedral Instability Governs Regret in Online Learning

Analyzes regret in online learning over combinatorial actions using convex relaxations, governed by polyhedral instability.

Ax Valentin Six, Frederik Panse, Mathis Fajeau, Lancelot Da Costa, Mridul Sharma, Alfonso Amayuelas, Tim Z. Xiao, David Hyland, Philipp Hennig, Bernhard Sch\"olkopf 5/14/2026

Learning POMDP World Models from Observations with Language-Model Priors

Learns POMDP world models from observation-action trajectories using language model priors for agents in partially observable environments.