Ax Yoav Gelberg, Yam Eitan, Michael Bronstein, Yarin Gal, Haggai Maron 5/8/2026

Training Transformers for KV Cache Compressibility

Research on training transformers to be more compressible for KV cache compression, addressing long-context language modeling bottlenecks.

Ax Licheng Pan, Haochen Yang, Haoxuan Li, Yunsheng Lu, Yongqi Tong, Yinuo Wang, Shijian Wang, Zhixuan Chu, Lei Shen, Yuan Lu, Hao Wang 5/8/2026

Optimal Transport for LLM Reward Modeling from Noisy Preference

Framework using optimal transport theory to train robust reward models for RLHF despite noisy preference data.

Ax Yingxu Wang, Kunyu Zhang, Yanwu Yang, Thomas Wolfers, Yujie Wu, Siyang Gao, Nan Yin 5/8/2026

When Brain Networks Travel: Learning Beyond Site

Cross-site fMRI graph learning method addressing out-of-distribution generalization across brain imaging sites via transient neurodynamics.

Ax Maxim Fishman, Brian Chmiel, Ron Banner, Daniel Soudry, Boris Ginsburg 5/8/2026

Normalized Architectures are Natively 4-Bit

nGPT architecture with unit hypersphere constraints enables stable 4-bit LLM training without random transforms or per-tensor scaling.

Ax Mikalai Korbit, Mario Zanon 5/8/2026

Fast Gauss-Newton for Multiclass Cross-Entropy

Fast Gauss-Newton decomposes multiclass softmax cross-entropy curvature into true-vs-rest and within-competitor terms for scaled optimization.

Ax Samir Darouich, Vinh Tong, Llu\'is Pastor-P\'erez, Tanja Bien, Loay Mualem, Mathias Niepert 5/8/2026

SymDrift: One-Shot Generative Modeling under Symmetries

SymDrift enables one-shot generative modeling of physical systems via equivariant diffusion models that preserve global symmetries like rotations.

Ax Jan von Pichowski, Al\v{z}beta Hrabo\v{s}ov\'a, Ingo Scholtes, Christopher Bl\"ocker 5/8/2026

The Role of Node Features in Graph Pooling

Study analyzing how node features and graph topology interact in graph pooling for classification tasks, examining conditions for effective pooling operators.