Ax Divya Shyamal, Marta Kne\v{z}evi\'c, Lan Tran, Chanakya Ekbote, Vijay Lingam, Paul Pu Liang 4/22/2026

SCATR: Simple Calibrated Test-Time Ranking

SCATR method for test-time scaling in LLMs using efficient ranking scorers as alternative to expensive process reward models.

Ax Tiankai Yang, Yi Nian, Xinyuan Li, Ruiyao Xu, Kaize Ding, Yue Zhao 4/22/2026

Cat-DPO: Category-Adaptive Safety Alignment

Category-adaptive safety alignment method for LLMs that addresses safety across different harm categories rather than uniform safety scoring.

Ax Liubomyr Horbatko 4/22/2026

Sessa: Selective State Space Attention

Selective State Space Attention architecture combining transformer attention with state-space models for long-context sequences.

Ax Yuyuan Chen, Shiyi Wang, Peter Potaptchik, Jaeyeon Kim, Michael S. Albergo 4/22/2026

Discrete Tilt Matching

Discrete Tilt Matching method for fine-tuning masked diffusion LLMs using likelihood-free RL objectives.

Ax Benjamin K. Johnson, Thomas Goralski, Ayush Semwal, Hui Shen, H. Josh Jang 4/22/2026

Streaming Structured Inference with Flash-SemiCRF

Technical paper on Semi-Markov CRFs for efficient streaming structured sequence inference without materializing large tensors.

Ax Priyam Dey, Aditya Sahdev, Sunny Bhati, Konda Reddy Mopuri, R. Venkatesh Babu 4/22/2026

Rethinking Dataset Distillation: Hard Truths about Soft Labels

Analysis of dataset distillation showing soft labels mask underperformance versus random baselines; proposes hard-label evaluation for reliable coreset comparison.

Ax Pourya Shamsolmoali, Masoumeh Zareapoor, Eric Granger, William A. P. Smith, Yue Lu 4/22/2026

Task Switching Without Forgetting via Proximal Decoupling

Proximal decoupling approach for continual learning addressing task switching without catastrophic forgetting through explicit separation of learning and retention signals.

Ax Chih-Yu Chang, Qiyuan Chen, Tianhan Gao, David Fenning, Chinedum Okwudire, Neil Dasgupta, Wei Lu, Raed Al Kontar 4/22/2026

Collaborative Contextual Bayesian Optimization

Collaborative contextual Bayesian optimization extending BO to context-specific optimal design problems requiring mapping from context space to solutions.

Ax Nathaniel S. Woodward, Zhiqi Gao, Yurii Kvasiuk, Kendrick M. Smith, Frederic Sala, Moritz M\"unchmeyer 4/22/2026

Fine-Tuning Small Reasoning Models for Quantum Field Theory

Fine-tuning study of small 7B reasoning models on theoretical physics with analysis of domain-specific reasoning ability development in LLMs.

Ax Arsalan Sharifnassab, Mohamed Elsayed, Kris De Asis, A. Rupam Mahmood, Richard S. Sutton 4/22/2026

Intentional Updates for Streaming Reinforcement Learning

Proposes intentional updates for stable streaming reinforcement learning by solving for step sizes matching intended outcome changes.

Ax Qingyang Zhang, Xinke Kong, Haitao Wu, Qinghua Hu, Minghao Wu, Baosong Yang, Yu Cheng, Yun Luo, Ganqu Cui, Changqing Zhang 4/22/2026

TEMPO: Scaling Test-time Training for Large Reasoning Models

TEMPO scales test-time training for large reasoning models by addressing reward signal drift with external calibration mechanisms.