Ax Thibault Pautrel, Florent Bouchard, Ammar Mian, Guillaume Ginolhac 4/27/2026

FedSPDnet: Geometry-Aware Federated Deep Learning with SPDnet

Research paper on federated learning frameworks for SPDnet models operating on symmetric positive definite matrices with geometry-preserving aggregation strategies.

Ax Coenraad Mouton, Randle Rabe, Niklas C. Koser, Nicolai Krekiehn, Christopher Hansen, Jan-Bernd H\"ovener, Claus-C. Gl\"uer 4/27/2026

Useful nonrobust features are ubiquitous in biomedical images

Studies nonrobust predictive features in deep medical imaging models. Shows networks learn adversarially vulnerable patterns useful in-distribution.

Ax Zehua Pei, Ying Zhang, Hui-Ling Zhen, Tao Yuan, Xianzhi Yu, Zhenhua Dong, Sinno Jialin Pan, Mingxuan Yuan, Bei Yu 4/27/2026

PreMoE: Proactive Inference for Efficient Mixture-of-Experts

PreMoE: training-free framework to compile sparse Mixture-of-Experts variants for deployment-specific optimization using predicted expert utility.

Ax Omkar Tupe, Max Hartman, Lav R. Varshney, Saurav Prakash 4/27/2026

Federated Nonlinear System Identification

Federated learning framework for nonlinear system identification with theoretical convergence guarantees improving with client count.

Ax Peng Chen, Jiaji Zhang, Hailiang Zhao, Yirong Zhang, Shenyao Chen, Jiahong Yu, Xueyan Tang, Yixuan Wang, Hao Li, Jianping Zou, Gang Xiong, Kingsum Chow, Shuibing He, Shuiguang Deng 4/27/2026

Toward Robust and Efficient ML-Based GPU Caching for Modern Inference

Learning-augmented caching algorithm for GPU inference that combines ML predictors with robustness guarantees against prediction errors.

Ax Rebonto Haque, Oliver M. Turnbull, Anisha Parsan, Nithin Parsan, John J. Yang, Anna L. Beukenhorst, Charlotte M. Deane 4/27/2026

Mechanistic Interpretability of Antibody Language Models Using SAEs

Mechanistic interpretability of antibody language models using sparse autoencoders for feature discovery and steering in protein sequence generation.

Ax Marcel Meyer, Sascha Kaltenpoth, Henrik Albers, Kevin Zalipski, Oliver M\"uller 4/27/2026

TS-Arena -- A Live Forecast Pre-Registration Platform

TS-Arena live forecasting platform for evaluating time series foundation models on unknown future data, addressing train-test contamination issues.

Ax Deming Chen, Vijay Ganesh, Weikai Li, Yingyan Celine Lin, Yong Liu, Subhasish Mitra, David Z. Pan, Ruchir Puri, Jason Cong, Yizhou Sun 4/27/2026

Report for NSF Workshop on AI for Electronic Design Automation

NSF workshop report on AI for Electronic Design Automation discussing LLMs, GNNs, RL, and neurosymbolic methods for EDA automation.

Ax Amin Oji, Paul Fieguth 4/27/2026

Joint Embedding Variational Bayes

Variational Joint Embedding framework for non-contrastive self-supervised learning using symmetric conditional ELBO on paired encoder embeddings.

Ax Noor Islam S. Mohammad, Md Muntaqim Meherab 4/27/2026

Regularized Meta-Learning for Improved Generalization

Regularized meta-learning framework addressing redundancy, multicollinearity, and overfitting in deep ensemble methods through four-stage projection pipeline.

Ax Aleena Siji, Amir Mohammad Karimi Mamaghan, Ferdinand Kapl, Tobias H\"oppe, Emmanouil Angelis, Andrea Dittadi, Maurice Brenner, Michael Heinzinger, Karl Henrik Johansson, Kaitlin Maile, Johannes von Oswald, Stefan Bauer 4/27/2026

From Words to Amino Acids: Does the Curse of Depth Persist?

Investigates curse of depth in protein language models, comparing depth scaling effects between protein and natural language transformers.

Ax Md Muntaqim Meherab, Noor Islam S. Mohammad, Faiza Feroz 4/27/2026

Causal Concept Graphs in LLM Latent Space for Stepwise Reasoning

Causal Concept Graphs combine sparse autoencoders with differentiable structure learning to capture causal dependencies between interpretable latent features for multi-step LLM reasoning.

Ax Hongtao Xu, Jianchao Tan, Yuxuan Hu, Pengju Lu, Hongyu Wang, Pingwei Sun, Yerui Sun, Yuchen Xie, Xunliang Cai, Mingzhen Li, Weile Jia 4/27/2026

SparseBalance: Load-Balanced Long Context Training with Dynamic Sparse Attention

SparseBalance addresses load balancing challenges in distributed sparse attention training for long-context LLMs by co-optimizing sequence length heterogeneity and sparsity sensitivity.

Ax Zhaokun Wang, Jinyu Guo, Jingwen Pu, Hongli Pu, Meng Yang, Xunlei Chen, Jie Ou, Wenyi Li, Guangchun Luo, Wenhong Tian 4/27/2026

CAP: Controllable Alignment Prompting for Unlearning in LLMs

CAP method enables selective knowledge unlearning in closed-source LLMs through controllable alignment prompting without model weight access.

Ax Sajad Movahedi, Timur Carstensen, Arshia Afzal, Frank Hutter, Antonio Orvieto, Volkan Cevher 4/27/2026

Selective Rotary Position Embedding

Selective RoPE: position encoding combining fixed-angle rotations from RoPE with input-dependent selective gating for improved language modeling.

Ax Sourav Saha, Mandar Mitra, Aditya Dutta 4/27/2026

LLMs as Assessors: Right for the Right Reason?

Study of whether LLMs as relevance assessors in IR tasks provide correct reasoning, extending research on LLMs as judges for output evaluation.

Ax John Kirchenbauer, Abhimanyu Hans, Brian Bartoldson, Micah Goldblum, Ashwinee Panda, Tom Goldstein 4/27/2026

Multi-Token Prediction via Self-Distillation

Self-distillation approach to convert pretrained autoregressive language models to multi-token prediction for faster inference without auxiliary models.