Ax Andrew Sellergren, Chufan Gao, Fereshteh Mahvar, Timo Kohlberger, Fayaz Jamil, Madeleine Traverse, Alberto Tono, Bashir Sadjad, Lin Yang, Charles Lau, Liron Yatziv, Tiffany Chen, Bram Sterling, Kenneth Philbrick, Richa Tiwari, Yun Liu, Madhuram Jajoo, Chandrashekar Sankarapu, Swapnil Vispute, Harshad Purandare, Abhishek Bijay Mishra, Sam Schmidgall, Tao Tu, Anil Palepu, Chunjong Park, Tim Strother, Rahul Thapa, Yong Cheng, Preeti Singh, Kat Black, Yossi Matias, Katherine Chou, Avinatan Hassidim, Kavi Goel, Joelle Barral, Tris Warkentin, Shravya Shetty, Dale Webster, Sunny Virmani, David F. Steiner, Can Kirmizibayrak, Daniel Golden 5/6/2026

MedGemma 1.5 Technical Report

Technical report on MedGemma 1.5 4B, a medical-specialized LLM adding imaging, anatomical localization, and medical document understanding capabilities.

Ax Bowen Ye, Rang Li, Qibin Yang, Yuanxin Liu, Linli Yao, Hanglong Lv, Zhihui Xie, Chenxin An, Lei Li, Lingpeng Kong, Qi Liu, Zhifang Sui, Tong Yang 5/6/2026

Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents

Introduces Claw-Eval, an end-to-end evaluation suite for autonomous LLM agents with 300 human-verified tasks addressing safety, robustness, and modality coverage.

Ax Ha Thanh Nguyen, Wachara Fungwacharakorn, Sabine Wehnert, May Myo Zin, Yuntao Kong, Jieying Xue, Micha{\l} Araszkiewicz, Randy Goebel, Ken Satoh 5/6/2026

GDPR Auto-Formalization with AI Agents and Human Verification

Arxiv paper on automatic GDPR formalization using multi-agent LLM workflow with role-specialized components and human-in-the-loop verification.

Ax Ifdita Hasan Orney, Jubayer Ibn Hamid, Shreya S Ramanujam, Shirley Wu, Hengyuan Hu, Noah Goodman, Dorsa Sadigh, Chelsea Finn 5/6/2026

Poly-EPO: Training Exploratory Reasoning Models

Arxiv paper on Poly-EPO framework for post-training language models to encourage optimistic exploration and balance exploration-exploitation trade-offs.

Ax Haebin Seong, Li Yin, Haoran Zhang, Zhan Shi 5/6/2026

The Last Harness You'll Ever Build

Arxiv paper on universal harness framework for AI agents navigating complex domain-specific workflows without painstaking task-specific engineering.

Ax Shuzheng Si, Haozhe Zhao, Yu Lei, Qingyi Wang, Dingwei Chen, Zhitong Wang, Zhenhailong Wang, Kangyang Luo, Zheng Wang, Gang Chen, Fanchao Qi, Minjia Zhang, Maosong Sun 5/6/2026

From Context to Skills: Can Language Models Learn from Context Skillfully?

Arxiv paper on context learning in language models via inference-time skill augmentation for reasoning over complex contexts exceeding parametric knowledge.

Ax Zhensu Sun, Haotian Zhu, Bowen Xu, Xiaoning Du, Li Li, David Lo 5/6/2026

Towards Agentic Runtime Healing

Paper on using LLMs for automated runtime healing in self-healing systems, replacing predefined rules with adaptive error recovery.

Ax Xuan Shen, Yizhou Wang, Yufa Zhou, Xiangxi Shi, Pu Zhao, Yanzhi Wang, Jiuxiang Gu 5/6/2026

Efficient Reasoning with Hidden Thinking

Research on Heima framework that compresses chain-of-thought reasoning in MLLMs into abstract thinking tokens for efficiency.

Ax Alexandra Bazarova, Aleksandr Yugay, Andrey Shulga, Alina Ermilova, Andrei Volodichev, Konstantin Polev, Julia Belikova, Rauf Parchiev, Dmitry Simakov, Maxim Savchenko, Andrey Savchenko, Serguei Barannikov, Alexey Zaytsev 5/6/2026

Hallucination Detection in LLMs with Topological Divergence on Attention Graphs

TOHA detector for identifying LLM hallucinations in RAG systems by analyzing topological divergence patterns in attention graph structures.

Ax Jiongli Zhu, Yue Wang, Bailu Ding, Philip A. Bernstein, Vivek Narasayya, Surajit Chaudhuri 5/6/2026

MINT: Multi-Vector Search Index Tuning

MINT framework for tuning index selection strategies in multi-vector databases to optimize performance across multiple feature dimensions.