Ax Hao Wang, Limeng Qiao, Chi Zhang, Lin Ma, Guanglu Wan, Xiangyuan Lan, Xiaodan Liang 5/6/2026

X2SAM: Any Segmentation in Images and Videos

X2SAM integrates MLLMs with foundation segmentation models enabling pixel-level perception from conversational instructions across images and videos.

Ax Md. Enamul Hoq, Wataru Uegami, Saghir Alfasly, Ghazal Alabtah, Sahar Rahimi Malakshan, Armita Kazemi, Alex T. Schmitgen, Fred Prior, H. R. Tizhoosh 5/6/2026

Retrieval-Guided Generation for Safer Histopathology Image Captioning

Retrieval-guided generation approach for medical image captioning reducing hallucinations and factual inconsistency in histopathology image descriptions.

Ax Anirudh Iyengar Kaniyar Narayana Iyengar, Tampu Ravi Kumar, Manan Suri, Raviteja Bommireddy, Dinesh Manocha, Puneet Mathur, Vivek Gupta 5/6/2026

DIAGRAMS: A Review Framework for Reasoning-Level Attribution in Diagram QA

DIAGRAMS: review framework for creating reasoning-level attribution in diagram QA, providing structured evidence annotation across visual diagrams and infographics.

Ax Daniel Song, Peter Ney, Cristina Menghini, Faizan Ahmad, Aidan Boyd, Nathaniel Li, Ziwen Han, Jean-Christophe Testud, Saisuke Okabayashi, Maeve Ryan, Jinpeng Miao, Hamza Kwisaba, Felix Binder, Spencer Whitman, Jim Gust, Esteban Arcaute, Dhaval Kapil, Jacob Kahn, Ayaz Minhas, Tristan Goodman, Lauren Deason, Alexander Vaughan, Shengjia Zhao, Summer Yue 5/6/2026

Code World Model Preparedness Report

Code World Model preparedness report: Meta's code generation and reasoning model assessed for frontier AI risks; found no additional catastrophic risks beyond current ecosystem.

Ax Caleb Talley, Vedant Tibrewal, Seun Adekunle, Weiwen Dong, Xinyu Wu, Fariha Sheikh 5/6/2026

Multi-Perspective Transformers in ARC-AGI-2 Challenge

Approach using multi-perspective transformers with test-time training and products of experts on ARC-AGI-2 visual reasoning benchmark.

Ax Tam Nguyen, Tu Anh Nguyen, Sina Alemohammad, Richard G. Baraniuk 5/6/2026

Minimizing Collateral Damage in Activation Steering

Method to minimize unintended side effects (collateral damage) when using activation steering to control LLM behavior through internal representation intervention.