Ax Markus Knauer, Edoardo Fiorini, Maximilian M\"uhlbauer, Stefan Schneyer, Promwat Angsuratanawech, Florian Samuel Lay, Timo Bachmann, Samuel Bustamante, Korbinian Nottensteiner, Freek Stulp, Alin Albu-Sch\"affer, Jo\~ao Silv\'erio, Thomas Eiband 4/23/2026

MOMO: A framework for seamless physical, verbal, and graphical robot skill learning and adaptation

Interactive framework for robot skill adaptation using kinesthetic, natural language, and graphical modalities for non-expert users.

Ax Nicola Bariletto, Huy Nguyen, Nhat Ho, Alessandro Rinaldo 4/23/2026

On Bayesian Softmax-Gated Mixture-of-Experts Models

Theoretical study of Bayesian mixture-of-experts models with softmax gating mechanisms, exploring their probabilistic properties and learning dynamics.

Ax Mahmoud Abdelmoneum, Pierfrancesco Beneventano, Tomaso Poggio 4/23/2026

pAI/MSc: ML Theory Research with Humans on the Loop

pAI/MSc: Open-source multi-agent system for automating academic ML research workflows from hypothesis to manuscript draft with human guidance.

Ax Nils Graef, Filip Makraduli, Andrew Wasielewski, Matthew Clapp 4/23/2026

FlashNorm: Fast Normalization for Transformers

Optimizes RMSNorm computation in LLMs by eliminating normalization weights through matrix folding for parallel execution.

Ax Binchi Zhang, Zihan Chen, Cong Shen, Jundong Li 4/23/2026

Verification of Machine Unlearning is Fragile

Demonstrates vulnerabilities in machine unlearning verification strategies used to validate data removal from models.

Ax Saurabh Singh, Dmitry Lagun 4/23/2026

Latent Stochastic Interpolants

Extends stochastic interpolants framework to latent space with joint optimization of encoder-decoder for generative modeling.

Ax Debanjan Dutta, Anish Chakrabarty, Faizanuddin Ansari, Swagatam Das 4/23/2026

On the Existence of Universal Simulators of Attention

Theoretical analysis of transformer expressivity and learnability through the lens of universal simulators of attention mechanisms.

Ax Lingyu Jiang, Dengzhe Hou, Yuping Wang, Yao Su, Shuo Xing, Wenjing Chen, Xin Zhang, Zhengzhong Tu, Ziming Zhang, Fangzhou Lin, Michael Zielewski, Kazunori D Yamada 4/23/2026

KANMixer: a minimal KAN-centered mixer for long-term time series forecasting

KANMixer architecture using Kolmogorov-Arnold Networks for long-term time series forecasting with improved expressivity over MLP and Transformer baselines.

Ax Feiyang Wu, Ye Zhao, Anqi Wu 4/23/2026

Distributional Inverse Reinforcement Learning

Distributional framework for offline inverse RL capturing richer expert behavior by modeling uncertainty over rewards and return distributions.

Ax Rishiraj Saha Roy, Chris Hinze, Luzian Hahn, Fabian Kuech 4/23/2026

CEDAR: Context Engineering for Agentic Data Science

CEDAR demonstrates automated data science task solving using LLM agents via context engineering, addressing task complexity, data size, and computational constraints.

Ax Mikael M{\o}ller H{\o}gsgaard, Chirag Pabbaraju 4/23/2026

Agnostic Language Identification and Generation

Theoretical analysis of language identification and generation without realizability assumptions, establishing statistical rates under relaxed conditions.

Ax Rong Fu, Muge Qi, Yang Li, Yabin Jin, Jiekai Wu, Jiaxuan Lu, Chunlei Meng, Youjin Wang, Zeli Su, Juntao Gao, Li Bao, Qi Zhao, Wei Luo, Simon Fong 4/23/2026

SwiftRepertoire: Few-Shot Immune-Signature Synthesis via Dynamic Kernel Codes

Few-shot learning framework for T cell receptor repertoire analysis using dynamic kernel codes and prototype-based parameterizations for disease detection.