Ax Ryan Wei Heng Quek, Sanghyuk Lee, Alfred Wei Lun Leong, Arun Verma, Alok Prakash, Nancy F. Chen, Bryan Kian Hsiang Low, Daniela Rus, Armando Solar-Lezama 5/15/2026

MeMo: Memory as a Model

MeMo framework encodes new knowledge into a dedicated memory model while keeping the LLM frozen for efficient knowledge updates.

Ax Michael Theologitis, Vasilis Samoladas, Antonios Deligiannakis 5/15/2026

Communication-Efficient Federated Fine-Tuning

Communication-efficient federated fine-tuning of language models with parameter compression for distributed learning scenarios.

Ax Alex Chen, Renato Geh, Aditya Grover, Guy Van den Broeck, Daniel Israel 5/15/2026

The Pitfalls of KV Cache Compression

Study identifying pitfalls in KV cache compression for LLMs in realistic multi-instruction scenarios with practical implications.

Ax Yihong Wu, Liheng Ma, Lei Ding, Muzhi Li, Xinyu Wang, Kejia Chen, Zhan Su, Zhanguang Zhang, Chenyang Huang, Yingxue Zhang, Mark Coates, Jian-Yun Nie 5/15/2026

It Takes Two: Your GRPO Is Secretly DPO

Analysis showing GRPO reinforcement learning algorithm for LLM post-training is equivalent to DPO with group-level baselines.

Ax Ning Yang, Hengyu Zhong, Haijun Zhang, Randall Berry 5/15/2026

Vision-LLMs for Spatiotemporal Traffic Forecasting

Vision-LLM approach for spatiotemporal traffic forecasting combining visual understanding of grid-based traffic data with language model capabilities.

Ax Robert Joseph George, Carson Eisenach, Udaya Ghai, Dominique Perrault-Joncas, Anima Anandkumar, Dean Foster 5/15/2026

BRIDGE: Building Representations In Domain Guided Program Synthesis

BRIDGE framework for structured prompting of LLMs to generate code with formal verification in proof assistants like Lean, handling multiple coupled domains.

Ax Jingkun Liu, Yisong Yue, Max Welling, Yue Song 5/15/2026

Krause Synchronization Transformers

Krause Attention mechanism addressing representation collapse and attention sink phenomena in transformers through principled bounded-confidence dynamics.

Ax Pascal Jr Tikeng Notsawo, Guillaume Dumas, Guillaume Rabusseau 5/15/2026

Grokking Finite-Dimensional Algebra

Study of grokking phenomenon (sudden generalization) in neural networks learning finite-dimensional algebra operations, extending prior work on group operations.

Ax Th\'eo Vincent, Kevin Gerhardt, Yogesh Tripathi, Habib Maraqten, Adam White, Martha White, Jan Peters, Carlo D'Eramo 5/15/2026

Gradient Iterated Temporal-Difference Learning

Gradient Iterated Temporal-Difference Learning addresses divergence issues in TD learning with semi-gradient updates.