Ax Andres Potapczynski, Ravi Kiran Selvam, Tatiana Konstantinova, Shankar Ramasubramanian, Malcolm Wolff, Kin G. Olivares, Ruijun Ma, Mengfei Cao, Michael W. Mahoney, Andrew Gordon Wilson, Boris N. Oreshkin, Dmitry Efimov 3/18/2026

Time-Aware Prior Fitted Networks for Zero-Shot Forecasting with Exogenous Variables

Zero-shot forecasting method for time series with exogenous variables using prior-fitted networks.

Ax Hanxian Huang, Igor Fedorov, Andrey Gromov, Bernard Beckerman, Naveen Suda, David Eriksson, Maximilian Balandat, Rylan Conway, Patrick Huber, Chinnadhurai Sankar, Ayushi Dalmia, Zechun Liu, Lemeng Wu, Tarek Elgamal, Adithya Sagar, Vikas Chandra, Raghuraman Krishnamoorthi 3/18/2026

MobileLLM-Flash: Latency-Guided On-Device LLM Design for Industry Scale

Hardware-in-the-loop architecture search methodology for designing efficient on-device LLMs with real-time latency constraints for mobile deployment.

Ax Swadesh Jana, Cansu Sancaktar, Tom\'a\v{s} Dani\v{s}, Georg Martius, Antonio Orvieto, Pavel Kolev 3/18/2026

GASP: Guided Asymmetric Self-Play For Coding LLMs

Proposes guided asymmetric self-play method for post-training coding LLMs with better problem selection to improve model capabilities.

Ax Xiaolong Han, Ferrante Neri, Zijian Jiang, Fang Wu, Yanfang Ye, Lu Yin, Zehong Wang 3/18/2026

W2T: LoRA Weights Already Know What They Can Do

Analyzes whether LoRA checkpoint weights encode task performance information readable without running the base model, enabling efficient adapter analysis.

Ax Christina Baek, Ricardo Pio Monti, David Schwab, Amro Abbas, Rishabh Adiga, Cody Blakeney, Maximilian B\"other, Paul Burstein, Aldo Gael Carranza, Alvin Deng, Parth Doshi, Vineeth Dorna, Alex Fang, Tony Jiang, Siddharth Joshi, Brett W. Larsen, Jason Chan Lee, Katherine L. Mentzer, Luke Merrick, Haakon Mongstad, Fan Pan, Anshuman Suri, Darren Teh, Jason Telanoff, Jack Urbanek, Zhengping Wang, Josh Wills, Haoli Yin, Aditi Raghunathan, J. Zico Kolter, Bogdan Gaza, Ari Morcos, Matthew Leavitt, Pratyush Maini 3/18/2026

The Finetuner's Fallacy: When to Pretrain with Your Finetuning Data

Study of specialized pretraining strategy using domain data during pretraining to improve finetuning performance and reduce forgetting.

Ax Hoang Phan, Quang H. Nguyen, Hung T. Q. Le, Xiusi Chen, Heng Ji, Khoa D. Doan 3/18/2026

Decoding the Critique Mechanism in Large Reasoning Models

Study of how large reasoning models use backtracking and self-verification to detect and correct errors in complex logical reasoning tasks.

Ax Laurent Cheret, Vincent L\'etourneau, Isar Nejadgholi, Chris Drummond, Hussein Al Osman, Maia Fraser 3/18/2026

Manifold-Matching Autoencoders

Unsupervised autoencoder regularization by aligning pairwise distances between latent and input spaces on learned manifolds.

Ax Hangting Ye, Peng Wang, Wei Fan, Xiaozhuang Song, He Zhao, Dandan Gun, Yi Chang 3/18/2026

Deep Tabular Representation Corrector

Deep learning methods for tabular data using representation correction to improve on in-learning and pre-learning paradigms.

Ax Gregor Kornhardt, Jannis Chemseddine, Christian Wald, Gabriele Steidl 3/18/2026

Self-Aware Markov Models for Discrete Reasoning

Method for discrete reasoning using self-aware Markov models that correct errors in masked diffusion models through adaptive denoising.