Ax Dario Shariatian, Alain Durmus, Umut Simsekli, Stefano Peluchetti 5/14/2026

Latent-Augmented Discrete Diffusion Models

Latent-Augmented Discrete Diffusion: learnable auxiliary channels for improved few-step language generation.

Ax George Yakushev, Nataliia Babina, Masoud Vahid Dastgerdi, Vyacheslav Zhdanovskiy, Denis Kuznedelev, Alina Shutova, Max Ryabinin 5/14/2026

Asynchronous Reasoning: Training-Free Interactive Thinking LLMs

Training-free asynchronous reasoning for LLM agents enabling real-time responses without sequential thinking bottlenecks.

Ax Yifan Zhang, Yifeng Liu, Mengdi Wang, Quanquan Gu 5/14/2026

Deep Delta Learning

Deep Delta Learning: transformer residual update rule enabling selective rewriting of content while preserving identity paths.

Ax Nicolas Zumarraga, Thomas Kaar, Ning Wang, William Tennien, Alpay Hasanli, Max Rosenblattl, Fan Wu, Kevin Riehl, Maxwell A. Xu, Markus Kreft, Kevin O'Sullivan, Elgar Fleisch, Paul Schmiedmayer, Robert Jakob, Patrick Langer 5/14/2026

TS-Haystack: A Multi-Task Retrieval Benchmark for Long-Context Time-Series Reasoning

TS-Haystack: benchmark for evaluating time-series language models on retrieval and reasoning over long temporal contexts across 10 tasks.

Ax Alex Morehead, Miruna Cretu, Antonia Panescu, Rishabh Anand, Maurice Weiler, Tynan Perez, Samuel Blau, Steven Farrell, Wahid Bhimji, Anubhav Jain, Hrushikesh Sahasrabuddhe, Pietro Lio, Tommi Jaakkola, Rafael Gomez-Bombarelli, Rex Ying, N. Benjamin Erichson, Michael W. Mahoney 5/14/2026

Zatom-1: Towards a Multimodal Foundation Model for 3D Molecules and Materials

Zatom-1: multimodal foundation model unifying generative and predictive learning for 3D molecules and materials across domains.

Ax Tomas Ruiz, Zhen Qin, Yifan Zhang, Xuyang Shen, Yiran Zhong, Mengdi Wang 5/14/2026

FlashSampling: Fast and Memory-Efficient Exact Sampling

FlashSampling fuses categorical sampling into LM-head matmul for fast exact sampling without materializing logits in large-vocabulary decoding.

Ax Ian Osband 5/14/2026

Delightful Distributed Policy Gradient

Research on distributed reinforcement learning addressing negative learning from high-surprisal data in stale or mismatched actor settings.