Ax Qiyuan Zhang, Junyi Zhou, Yufei Wang, Fuyuan Lyu, Yidong Ming, Can Xu, Qingfeng Sun, Kai Zheng, Peng Kang, Xue Liu, Chen Ma 3/4/2026

RubricBench: Aligning Model-Generated Rubrics with Human Standards

RubricBench benchmark for evaluating rubric-guided LLM reward models against human standards, addressing discriminative complexity in alignment evaluation.

Ax Robin Young 3/4/2026

What Is the Alignment Tax?

Geometric theory formalizing alignment tax as projection in representation space, deriving Pareto frontier for safety-capability tradeoffs in LLMs.

Ax Hongjin Qian, Ziyi Xia, Ze Liu, Jianlyu Chen, Kun Luo, Minghao Qin, Chaofan Li, Lei Xiong, Junwei Lan, Sen Wang, Zhengyang Liang, Yingxia Shao, Defu Lian, Zheng Liu 3/4/2026

DeepXiv-SDK: An Agentic Data Interface for Scientific Literature

SDK providing LLM-agents structured data access to scientific literature via agentic interface, reducing token consumption and improving retrieval efficiency.

Ax Zhonghang Li, Zongwei Li, Yuxuan Chen, Han Shi, Jiawei Li, Jierun Chen, Haoli Bai, Chao Huang 3/4/2026

FastCode: Fast and Cost-Efficient Code Understanding and Reasoning

FastCode system for efficient repository-scale code reasoning using selective context retrieval and compression for cost-effective LLM-based software engineering.

Ax Adrian Robert Minut, Hazem Dewidar, Iacopo Masi 3/4/2026

Spilled Energy in Large Language Models

Reinterprets LLM softmax as energy-based model to track 'energy spills' during decoding, correlating them with factual errors and biases.

Ax Jingxuan Fan, Yueying Li, Zhenting Qi, Dinghuai Zhang, Kiant\'e Brantley, Sham M. Kakade, Hanlin Zhang 3/4/2026

Scaling Reward Modeling without Human Supervision

Unsupervised reward modeling scaling via preference learning on web document prefixes/suffixes, reducing human annotation costs.

Ax Jiace Zhu, Wentao Chen, Qi Fan, Zhixing Ren, Junying Wu, Xing Zhe Chai, Chotiwit Rungrueangwutthinon, Yehan Ma, An Zou 3/4/2026

CUDABench: Benchmarking LLMs for Text-to-CUDA Generation

CUDABench benchmark for evaluating LLM text-to-CUDA code generation with performance assessment metrics for GPU kernels.

Ax Laziz U. Abdullaev, Noelle Y. L. Wong, Ryan T. Z. Lee, Shiqi Jiang, Khoi N. M. Nguyen, Tan M. Nguyen 3/4/2026

Concept Heterogeneity-aware Representation Steering

Method for steering LLM behavior via representation manipulation that accounts for heterogeneous concept encoding across embedding spaces.

Ax Andy Yang, Pascal Bergstr\"a{\ss}er, Georg Zetzsche, David Chiang, Anthony W. Lin 3/4/2026

Length Generalization Bounds for Transformers

Theoretical analysis of length generalization bounds for transformers on CRASP language class, addressing model generalization guarantees.

Ax Shadab Ahamed, Eshed Gal, Simon Ghyselincks, Md Shahriar Rahim Siddiqui, Moshe Eliasof, Eldad Haber 3/4/2026

Preconditioned Score and Flow Matching

Preconditioning techniques for flow matching and score-based diffusion to improve optimization by handling ill-conditioned covariance matrices.

Ax Satish Chandran, Nicolas Roque dos Santos, Yunshu Wu, Greg Ver Steeg, Evangelos Papalexakis 3/4/2026

Spectral Regularization for Diffusion Models

Introduces loss-level spectral regularization using Fourier and wavelet-domain losses to improve diffusion model training without architecture changes.