Ax Ricardo Euler, Pedro Maristany de las Casas, Ralf Bornd\"orfer 5/6/2026

Logic-Constrained Shortest Paths for Flight Planning

Algorithm for logic-constrained shortest path problem with satisfiability constraints applied to flight planning optimization. Specialized constraint solving research.

Ax Hyunji Min, Sangwon Jung, Junyoung Sung, Dosung Lee, Leekyeung Han, Paul Hongsuck Seo 5/6/2026

GOAT: A Training Framework for Goal-Oriented Agent with Tools

GOAT training framework for LLM agents with tool use, synthesizes API execution data from documentation for fine-tuning without human annotation.

Ax Vojtech Franc, Jakub Paplham 5/6/2026

Epistemic Reject Option Prediction

Reject option prediction method quantifying both epistemic and aleatoric uncertainty for high-stakes prediction applications.

Ax Weihao Bo, Shan Zhang, Yanpeng Sun, Jingjing Wu, Qunyi Xie, Xiao Tan, Kunbin Chen, Wei He, Xiaofan Li, Na Zhao, Jingdong Wang, Zechao Li 5/6/2026

Agentic Learner with Grow-and-Refine Multimodal Semantic Memory

Agentic learner with multimodal semantic memory that grows and refines domain knowledge across multiple modalities to avoid repeating mistakes.

Ax Ali Ezzat Shahroor, Mohamed Bayan Kmainasi, Abul Hasnat, Dimitar Dimitrov, Giovanni Da San Martino, Preslav Nakov, Firoj Alam 5/6/2026

MemeLens: Multilingual Multitask VLMs for Memes

MemeLens is a multilingual multitask Vision Language Model for meme understanding across hate detection, propaganda, and sentiment tasks.

Ax Ruijie Shi, Houbin Zhang, Yuecheng Han, Yuheng Wang, Jingru Fan, Runde Yang, Yufan Dang, Huatao Li, Dewen Liu, Yuan Cheng, Chen Qian 5/6/2026

AgentXRay: White-Boxing Agentic Systems via Workflow Reconstruction

Introduces Agentic Workflow Reconstruction (AWR) task to synthesize explicit, interpretable workflows from opaque LLM-based agentic systems.

Ax Varun Ursekar (Emily), Apaar Shanker (Emily), Veronica Chatrath (Emily), Yuan (Emily), Xue, Sam Denton 5/6/2026

VeRO: An Evaluation Harness for Agents to Optimize Agents

Introduces VeRO, an evaluation harness for coding agents optimizing other agents through iterative edit-execute-evaluate cycles.

Ax Andrew Sellergren, Chufan Gao, Fereshteh Mahvar, Timo Kohlberger, Fayaz Jamil, Madeleine Traverse, Alberto Tono, Bashir Sadjad, Lin Yang, Charles Lau, Liron Yatziv, Tiffany Chen, Bram Sterling, Kenneth Philbrick, Richa Tiwari, Yun Liu, Madhuram Jajoo, Chandrashekar Sankarapu, Swapnil Vispute, Harshad Purandare, Abhishek Bijay Mishra, Sam Schmidgall, Tao Tu, Anil Palepu, Chunjong Park, Tim Strother, Rahul Thapa, Yong Cheng, Preeti Singh, Kat Black, Yossi Matias, Katherine Chou, Avinatan Hassidim, Kavi Goel, Joelle Barral, Tris Warkentin, Shravya Shetty, Dale Webster, Sunny Virmani, David F. Steiner, Can Kirmizibayrak, Daniel Golden 5/6/2026

MedGemma 1.5 Technical Report

Technical report on MedGemma 1.5 4B, a medical-specialized LLM adding imaging, anatomical localization, and medical document understanding capabilities.

Ax Bowen Ye, Rang Li, Qibin Yang, Yuanxin Liu, Linli Yao, Hanglong Lv, Zhihui Xie, Chenxin An, Lei Li, Lingpeng Kong, Qi Liu, Zhifang Sui, Tong Yang 5/6/2026

Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents

Introduces Claw-Eval, an end-to-end evaluation suite for autonomous LLM agents with 300 human-verified tasks addressing safety, robustness, and modality coverage.

Ax Ha Thanh Nguyen, Wachara Fungwacharakorn, Sabine Wehnert, May Myo Zin, Yuntao Kong, Jieying Xue, Micha{\l} Araszkiewicz, Randy Goebel, Ken Satoh 5/6/2026

GDPR Auto-Formalization with AI Agents and Human Verification

Arxiv paper on automatic GDPR formalization using multi-agent LLM workflow with role-specialized components and human-in-the-loop verification.

Ax Ifdita Hasan Orney, Jubayer Ibn Hamid, Shreya S Ramanujam, Shirley Wu, Hengyuan Hu, Noah Goodman, Dorsa Sadigh, Chelsea Finn 5/6/2026

Poly-EPO: Training Exploratory Reasoning Models

Arxiv paper on Poly-EPO framework for post-training language models to encourage optimistic exploration and balance exploration-exploitation trade-offs.

Ax Haebin Seong, Li Yin, Haoran Zhang, Zhan Shi 5/6/2026

The Last Harness You'll Ever Build

Arxiv paper on universal harness framework for AI agents navigating complex domain-specific workflows without painstaking task-specific engineering.