Ax Zigeng Chen, Gongfan Fang, Xinyin Ma, Ruonan Yu, Xinchao Wang 4/10/2026

DMax: Aggressive Parallel Decoding for dLLMs

DMax enables efficient parallel decoding in diffusion language models through progressive self-refinement.

Ax Andrey Bocharnikov, Ivan Ermakov, Denis Kuznedelev, Vyacheslav Zhdanovskiy, Yegor Yershov 4/10/2026

KV Cache Offloading for Context-Intensive Tasks

KV cache offloading technique to reduce memory and latency overhead for long-context LLM inference.

Ax Xiangru Jian, Hao Xu, Wei Pang, Xinjian Zhao, Chengyu Tao, Qixin Zhang, Xikun Zhang, Chao Zhang, Guanzhi Deng, Alex Xue, Juan Du, Tianshu Yu, Garth Tarr, Linqi Song, Qiuzhuang Sun, Dacheng Tao 4/10/2026

FORGE:Fine-grained Multimodal Evaluation for Manufacturing Scenarios

Benchmark dataset and evaluation for multimodal LLMs in manufacturing scenarios.

Ax Mohamed Ehab (Faculty of Computer Science, October University for Modern Science & Arts, Giza, Egypt), Ali Hamdi (Faculty of Computer Science, October University for Modern Science & Arts, Giza, Egypt), Khaled Shaban (Department of Computer Science and Engineering, Qatar University, Doha, Qatar) 4/10/2026

CAMO: A Class-Aware Minority-Optimized Ensemble for Robust Language Model Evaluation on Imbalanced Data

CAMO is an ensemble technique for imbalanced text classification that optimizes minority class performance through hierarchical voting, confidence calibration, and uncertainty estimation.

Ax Mohammad Siavashi, Mariano Scazzariello, Gerald Q. Maguire Jr., Dejan Kosti\'c, Marco Chiesa 4/10/2026

Blink: CPU-Free LLM Inference by Delegating the Serving Stack to GPU and SmartNIC

Blink is an LLM serving architecture that removes the host CPU from the critical path by delegating orchestration and token control to GPU and SmartNIC, improving inference performance and datacenter resource utilization.

Ax Yuanjian Xu, Tianze Sun, Changwei Xu, XinLong Zhao, Jianing Hao, Ran Chen, Yang Liu, Ruijie Xu, Stephen Chen, Guang Zhang 4/10/2026

Rethinking Data Mixing from the Perspective of Large Language Models

Studies data mixing strategies for LLM training, questioning domain definitions, human-model alignment, and impact of domain weighting on generalization.

Ax Yanling Xiao, Huaibing Xie, Guoliang Zhao, Shihan Dou, Shaolei Wang, Yiting Liu, Nantao Zheng, Cheng Zhang, Pluto Zhou, Zhisong Zhang, Lemao Liu 4/10/2026

A Decomposition Perspective to Long-context Reasoning for LLMs

Decomposes long-context reasoning in LLMs into atomic skills, automatically identifying and improving fundamental capabilities for complex reasoning.

Ax Vladimir Zaigrajew, Micha{\l} Piechota, Gaspar Sekula, Przemys{\l}aw Biecek 4/10/2026

LINE: LLM-based Iterative Neuron Explanations for Vision Models

LINE uses LLMs iteratively to explain individual neuron concepts in vision models without predefined vocabularies, enabling interpretability of neural networks.