Ax Rui Xiao, Sanghwan Kim, Yongqin Xian, Zeynep Akata, Stephan Alaniz 3/19/2026

FINER: MLLMs Hallucinate under Fine-grained Negative Queries

Benchmark and analysis of hallucinations in multimodal LLMs with fine-grained negative queries covering multi-object, multi-attribute, and multi-relation scenarios.

Ax Yihong Chen, Quanming Yao 3/19/2026

Attention Sinks Induce Gradient Sinks

Analysis showing attention sinks in transformers induce gradient concentration during backpropagation under causal masking, affecting training dynamics.

Ax Luca Hinkamp, Simon Kl\"uttermann, Emmanuel M\"uller 3/19/2026

RangeAD: Fast On-Model Anomaly Detection

On-model anomaly detection method that leverages primary model representations to detect distributional shifts without separate AD models.

Ax Lintang Sutawika, Aditya Bharat Soni, Bharath Sriraam R R, Apurva Gandhi, Taha Yassine, Sanidhya Vijayvargiya, Yuchen Li, Xuhui Zhou, Yilin Zhang, Leander Melroy Maben, Graham Neubig 3/19/2026

CodeScout: An Effective Recipe for Reinforcement Learning of Code Search Agents

Reinforcement learning method for training code search agents to localize relevant files, classes, and functions in large repositories as prerequisite for coding tasks.

Ax Dharshan Kumaran, Arthur Conmy, Federico Barbero, Simon Osindero, Viorica Patraucean, Petar Velickovic 3/19/2026

How do LLMs Compute Verbal Confidence

arXiv paper proposing GeCO, time-unconditional flow matching framework for adaptive robotic control using diffusion models.

Ax Alexander D. Goldie, Zilin Wang, Adrian Hayler, Deepak Nathani, Edan Toledo, Ken Thampiratwong, Aleksandra Kalisz, Michael Beukman, Alistair Letcher, Shashank Reddy, Clarisse Wibault, Theo Wolf, Charles O'Neill, Uljad Berdica, Nicholas Roberts, Saeed Rahmani, Hannah Erlebach, Roberta Raileanu, Shimon Whiteson, Jakob N. Foerster 3/19/2026

Procedural Generation of Algorithm Discovery Tasks in Machine Learning

arXiv paper investigating how LLMs compute verbal confidence scores and whether they're generated just-in-time or cached during inference.

Ax Mohamed Eltahir, Ali Habibullah, Yazan Alshoibi, Lama Ayash, Tanveer Hussain, Naeemullah Khan 3/19/2026

VideoAtlas: Navigating Long-Form Video in Logarithmic Compute

Method for efficient long-form video processing with hierarchical grid representation enabling lossless, scalable video understanding.

Ax Jianrui Zhang, Yue Yang, Rohun Tripathi, Winson Han, Ranjay Krishna, Christopher Clark, Yong Jae Lee, Sangho Lee 3/19/2026

Unified Spatio-Temporal Token Scoring for Efficient Video VLMs

Token pruning approach for efficient video vision-language models addressing temporal redundancy in video-based downstream tasks.

Ax Oshadha Wijerathne (University of Moratuwa, Sri Lanka), Amandi Nimasha (University of Moratuwa, Sri Lanka), Dushan Fernando (University of Moratuwa, Sri Lanka), Nisansa de Silva (University of Moratuwa, Sri Lanka), Srinath Perera (WSO2 LLC) 3/19/2026

ScheduleMe: Multi-Agent Calendar Assistant

Multi-agent calendar assistant using graph-structured coordination with supervisory agent overseeing specialized task agents for natural language Google Calendar management.

Ax Dachuan Lin, Guobin Shen, Zihao Yang, Tianrong Liu, Dongcheng Zhao, Yi Zeng 3/19/2026

Efficient LLM Safety Evaluation through Multi-Agent Debate

Research on improving LLM safety evaluation using multi-agent debate with HAJailBench, a 11,100-sample human-annotated jailbreak benchmark across diverse attack methods.

Ax Sunghyun Wee, Suyoung Kim, Hyeonjin Kim, Kyomin Hwang, Nojun Kwak 3/19/2026

Safety-Preserving PTQ via Contrastive Alignment Loss

Safety-preserving post-training quantization via contrastive alignment loss to maintain behavioral safety during LLM model compression.