Ax Zijian Gao, Wangwang Jia, Xingxing Zhang, Pengfei Qian, Tao Sun, Bo Ding, Yong Dou, Huaimin Wang, Kele Xu 4/16/2026

MAny: Merge Anything for Multimodal Continual Instruction Tuning

MAny: Method for multimodal continual instruction tuning of MLLMs. Addresses catastrophic forgetting via parameter merging across perception and reasoning spaces.

Ax Gerg\H{o} Szalay, Gergely Zsolt Kov\'acs, S\'andor Teleki, Bal\'azs Pint\'er, Tibor Gregorics 4/16/2026

Neural architectures for resolving references in program code

Neural architectures for resolving code references and indirect indexing. Proposes seq2seq models for decompilation tasks with synthetic benchmarks.

Ax Yuanda Xu, Hejian Sang, Zhengze Zhou, Ran He, Zhipeng Wang, Alborz Geramifard 4/16/2026

TIP: Token Importance in On-Policy Distillation

Token Importance in on-policy knowledge distillation for LLMs. Identifies which token positions provide useful learning signals during student training on teacher supervision.

Ax Sumeet Ramesh Motwani, Daniel Nichols, Charles London, Peggy Li, Fabio Pizzati, Acer Blake, Hasan Hammoud, Tavish McDonald, Akshat Naik, Alesia Ivanova, Vignesh Baskaran, Ivan Laptev, Ruben Glatt, Tal Ben-Nun, Philip Torr, Natasha Jaques, Ameya Prabhu, Brian Bartoldson, Bhavya Kailkhura, Christian Schroeder de Witt 4/16/2026

LongCoT: Benchmarking Long-Horizon Chain-of-Thought Reasoning

LongCoT: benchmark of 2,500 expert-designed problems measuring long-horizon chain-of-thought reasoning across chemistry, math, CS, chess, and logic.

Ax Maksim Ivanov, Abhijay Rana, Gokul Prabhakaran 4/16/2026

Can Coding Agents Be General Agents?

Case study evaluating coding agents on business process automation tasks in ERP systems, identifying capability gaps beyond software engineering.

Ax Dikshant Kukreja (IIIT Delhi, India), Kshitij Sah (IIIT Delhi, India), Gautam Gupta (IIIT Delhi, India), Avinash Anand (Singapore Institute of Technology), Rajiv Ratn Shah (IIIT Delhi, India), Zhengkui Wang (Singapore Institute of Technology), Aik Beng Ng (NVIDIA), Erik Cambria (Nanyang Technological University) 4/16/2026

Better and Worse with Scale: How Contextual Entrainment Diverges with Model Size

Scaling laws for contextual entrainment showing larger language models simultaneously improve at ignoring false claims but worsen at ignoring irrelevant tokens.

Ax Hongyi Jin, Bohan Hou, Guanjie Wang, Ruihang Lai, Jinqi Chen, Zihao Ye, Yaxing Cai, Yixin Dong, Xinhao Cheng, Zhihao Zhang, Yilong Zhao, Yingyi Huang, Lijie Yang, Jinchen Jiang, Gabriele Oliaro, Jianan Ji, Xupeng Miao, Vinod Grover, Todd C. Mowry, Zhihao Jia, Tianqi Chen 4/16/2026

Event Tensor: A Unified Abstraction for Compiling Dynamic Megakernel

Event Tensor compiler framework unifying dynamic megakernel abstraction to improve LLM inference performance by eliminating kernel launch overhead and enabling inter-kernel parallelism.

Ax Joel Niklaus, Atsuki Yamaguchi, Michal \v{S}tef\'anik, Guilherme Penedo, Hynek Kydl\'i\v{c}ek, Elie Bakouch, Lewis Tunstall, Edward Emanuel Beeching, Thibaud Frere, Colin Raffel, Leandro von Werra, Thomas Wolf 4/16/2026

How Can We Synthesize High-Quality Pretraining Data? A Systematic Study of Prompt Design, Generator Model, and Source Data

Systematic study of synthetic data generation for LLM pretraining, testing rephrasing strategies, generator models, and source data across one trillion tokens to identify optimal design choices.

Ax Davyd Naveriani, Albert Zeyer, Ralf Schl\"uter, Hermann Ney 4/16/2026

Diffusion Language Models for Speech Recognition

Exploration of masked and uniform-state diffusion language models for speech recognition rescoring and ASR hypothesis improvement.

Ax Vansh Kapoor, Jayakrishnan Nair 4/16/2026

MDPs with a State Sensing Cost

Markov decision processes with state sensing costs, balancing optimal actions against sensing/communication/computation expenses in decision-making.

Ax Mingkuan Feng, Jinyang Wu, Siyuan Liu, Shuai Zhang, Ruihan Jin, Feihu Che, Pengpeng Shao, Zhengqi Wen, Jianhua Tao 4/16/2026

Two-Stage Regularization-Based Structured Pruning for LLMs

Two-stage regularization-based structured pruning method for reducing LLM parameters while minimizing knowledge loss and retraining requirements.

Ax William Anderson, Seung Whan Chung, Robert Stephany, Youngsoo Choi 4/16/2026

mLaSDI: Multi-stage latent space dynamics identification

Multi-stage latent space dynamics identification framework for solving PDEs via data-driven reduced-order models using autoencoders and ODEs.