Ax Murat Onur Yildirim, Elif Ceren Gok Yildirim, Joaquin Vanschoren 2/20/2026

Unlocking [CLS] Features for Continual Post-Training

Continual learning approach for foundation models addressing stability-plasticity trade-off during post-training on new classes/domains.

Ax Mark Lee, Chang Lan, Tom Gunter, John Peebles, Hanzhi Zhou, Kelvin Zou, Sneha Bangalore, Chung-Cheng Chiu, Nan Du, Xianzhi Du, Philipp Dufter, Ruixuan Hou, Haoshuo Huang, Dongseong Hwang, Xiang Kong, Jinhao Lei, Tao Lei, Meng Li, Li Li, Jiarui Lu, Zhiyun Lu, Yiping Ma, David Qiu, Vivek Rathod, Senyu Tong, Zhucheng Tu, Jianyu Wang, Yongqiang Wang, Zirui Wang, Floris Weers, Sam Wiseman, Guoli Yin, Bowen Zhang, Xiyou Zhou, Danyang Zhuo, Cheng Leong, Ruoming Pang 2/20/2026

AXLearn: Modular, Hardware-Agnostic Large Model Training

AXLearn production system for scalable hardware-agnostic training of large models with modular software architecture.

Ax Thibaud Gloaguen, Robin Staab, Nikola Jovanovi\'c, Martin Vechev 2/20/2026

Watermarking Diffusion Language Models

First watermarking method for diffusion language models that generate tokens non-sequentially, addressing unique DLM challenges.

Ax Ruchi Sandilya, Sumaira Perez, Charles Lynch, Lindsay Victoria, Benjamin Zebley, Derrick Matthew Buchanan, Mahendra T. Bhati, Nolan Williams, Timothy J. Spellman, Faith M. Gunning, Conor Liston, Logan Grosenick 2/20/2026

Contrastive Diffusion Alignment: Learning Structured Latents for Controllable Generation

Introduces ConDA, a contrastive learning layer for organizing diffusion model latent spaces to enable controllable generation.

Ax Yijun Ma, Zehong Wang, Weixiang Sun, Yanfang Ye 2/20/2026

Temporal Graph Pattern Machine

Temporal graph pattern machine for learning transferable representations in dynamic networks without restrictive assumptions.

Ax Beatrix M. G. Nielsen, Emanuele Marconato, Luigi Gresele, Andrea Dittadi, Simon Buchholz 2/20/2026

Logit Distance Bounds Representational Similarity

Analysis showing logit distance bounds representational similarity in discriminative models including autoregressive language models.

Ax Rachel Ma, Jingyi Qu, Andreea Bobu, Dylan Hadfield-Menell 2/20/2026

Goal Inference from Open-Ended Dialog

Framework for embodied AI agents to infer user goals from open-ended dialog using LLMs for efficient task accomplishment.

Ax Andr\'e Barreto, Vincent Dumoulin, Yiran Mao, Mark Rowland, Nicolas Perez-Nieves, Bobak Shahriari, Yann Dauphin, Doina Precup, Hugo Larochelle 2/20/2026

Capturing Individual Human Preferences with Reward Features

Learning user-specialized reward models for reinforcement learning from human feedback to capture individual preference disagreement.

Ax Mert Cemri, Nived Rajaraman, Rishabh Tiwari, Xiaoxuan Liu, Kurt Keutzer, Ion Stoica, Kannan Ramchandran, Ahmad Beirami, Ziteng Sun 2/20/2026

$\texttt{SPECS}$: Faster Test-Time Scaling through Speculative Drafts

SPECS: method for faster test-time scaling in LLMs through speculative drafts, balancing reasoning accuracy with user-facing latency.

Ax Yasaman Haghighi, Bastien van Delft, Mariam Hassan, Alexandre Alahi 2/20/2026

LayerSync: Self-aligning Intermediate Layers

LayerSync regularizes diffusion models using their own intermediate layer representations to improve generation quality and training efficiency.

Ax Marisa C. Peczuh, Nischal Ashok Kumar, Ryan Baker, Blair Lehman, Danielle Eisenberg, Caitlin Mills, Payu Wittawatolarn, Kushaan Naskar, Keerthi Chebrolu, Sudhip Nashi, Cadence Young, Brayden Liu, Sherry Lachman, Andrew Lan 2/20/2026

Toward LLM-Supported Automated Assessment of Critical Thinking Subskills

Uses LLMs for automated assessment of critical thinking skills in educational contexts, addressing evaluation of evidence and claim reliability.

Ax Ahmed Aboulfotouh, Hatem Abou-Zeid 2/20/2026

Multimodal Wireless Foundation Models

Extends wireless foundation models to accept multiple input modalities for improved task performance and adaptation across varying conditions.

Ax Mozes Jacobs, Thomas Fel, Richard Hakim, Alessandra Brondetta, Demba Ba, T. Andy Keller 2/20/2026

Block-Recurrent Dynamics in Vision Transformers

Introduces Block-Recurrent Hypothesis explaining Vision Transformer depth as block-recurrent computational flow for mechanistic interpretation.