Ax Parth Asawa, Alan Zhu, Abigail O'Neill, Matei Zaharia, Alexandros G. Dimakis, Joseph E. Gonzalez 5/18/2026

How to Train Your Advisor: Steering Black-Box LLMs with Advisor Models

Method to train small open-weight advisor models that generate dynamic prompts to improve black-box LLM performance. Demonstrates 27.4% improvement on GPT-5.2 tax tasks.

Ax Jamison Meindl, Yunsheng Tian, Tony Cui, Veronika Thost, Zhang-Wei Hong, Jie Chen, Wojciech Matusik, Mina Konakovi\'c Lukovi\'c 5/18/2026

SemanticOpt: Towards LLM-Based Semantic Black-Box Optimization

LLM-based semantic optimization approach for expensive black-box problems incorporating domain knowledge and heuristics.

Ax Yue Wang, Qizhou Wang, Zizhuo Zhang, Gang Niu, Bo Han, Masashi Sugiyama 5/18/2026

What Is Preference Optimization Doing, and Why?

Analysis of optimization dynamics comparing preference optimization methods like DPO and PPO for LLM alignment.

Ax Yixuan Even Xu, John Kirchenbauer, Yash Savani, Asher Trockman, Alexander Robey, Tom Goldstein, Fei Fang, J. Zico Kolter 5/18/2026

Antidistillation Fingerprinting

Fingerprinting method to detect when student LLMs trained on teacher model outputs via distillation

Ax Nick Alonso, Tomas Figliolia, Beren Millidge 5/18/2026

Online Vector Quantized Attention

Vector quantized attention layer balancing efficiency and performance for long context language models

Ax Junxiong Wang, Fengxiang Bie, Jisen Li, Zhongzhu Zhou, Zelei Shao, Yubo Wang, Yinghui Liu, Qingyang Wu, Avner May, Sri Yanamandra, Ce Zhang, Tri Dao, Percy Liang, Ben Athiwaratkun, Shuaiwen Leon Song, Chenfeng Xu, Xiaoxia Wu 5/18/2026

When RL Meets Adaptive Speculative Training: A Unified Training-Serving System

Unified system combining reinforcement learning with adaptive speculative decoding for efficient LLM serving

Ax Uzay Macar, Li Yang, Atticus Wang, Peter Wallich, Emmanuel Ameisen, Jack Lindsey 5/18/2026

Mechanisms of Introspective Awareness

Study of how LLMs detect and identify injected steering vectors in residual streams, revealing introspective awareness mechanisms.

Ax Zigeng Chen, Gongfan Fang, Xinyin Ma, Ruonan Yu, Xinchao Wang 5/18/2026

DMax: Aggressive Parallel Decoding for dLLMs

DMax paradigm for efficient parallel decoding in diffusion language models via progressive self-refinement, reducing error accumulation.

Ax Andrey Bocharnikov, Ivan Ermakov, Denis Kuznedelev, Vyacheslav Zhdanovskiy, Yegor Yershov 5/18/2026

KV Cache Offloading for Context-Intensive Tasks

Research on KV cache offloading to reduce memory and latency bottlenecks for long-context LLM inference tasks.

Ax Zhaokun Wang, Jinyu Guo, Jingwen Pu, Hongli Pu, Meng Yang, Xunlei Chen, Jie Ou, Wenyi Li, Guangchun Luo, Wenhong Tian 5/18/2026

CAP: Controllable Alignment Prompting for Unlearning in LLMs

Research on unlearning sensitive information from LLMs via controllable alignment prompting without modifying weights, addressing regulatory compliance.

Ax Zhaoyuan Su, Olatunji Ruwase, Karthik Ganesan, Aurick Qiao, Samyam Rajbhandari, Juncheng Yang, Yue Cheng, Yuxiong He 5/18/2026

MoE-Prefill: Zero Redundancy Overheads in MoE Prefill Serving

System for serving mixture-of-experts LLMs on prefill-only workloads without redundant distributed execution overhead.

Ax Thomas Walker, T. Mitchell Roddenberry, Ahmed Imtiaz Humayun, Randall Balestriero, Richard Baraniuk 5/18/2026

The Geometric Structure of Models Learning Sparse Data

Theoretical analysis of how models learn sparse data regimes where the manifold hypothesis does not apply.

Ax Christopher Lohse, Anish Dhir, Amadou Ba, Bradley Eck, Marco Ruffini, Jonas Wahl 5/18/2026

PRIM: Meta-Learned Bayesian Root Cause Analysis

Meta-learning approach framing root cause analysis in complex systems as Bayesian inference over synthetic causal model priors.