Ax Yangfu Li, Yuning Gong, Hongjian Zhan, Teng Li, Yuanhuiyi Lyu, Tianyi Chen, Qi Liu, Ziyuan Huang, Zhihang Zhong, Dandan Zheng, Yue Lu 5/6/2026

Perceptual Flow Network for Visually Grounded Reasoning

Vision-language model for grounded reasoning using perceptual flow networks to reduce hallucination and language bias.

Ax Frederic Grabowski, Jacek Szczerbi\'nski, Maciej Ja\'skowski, Kalina Jasi\'nska-Kobus, Pawe{\l} D\k{a}browski-Tuma\'nski, Tomasz Jetka, Bartosz Topolski 5/6/2026

Bolek: A Multimodal Language Model for Molecular Reasoning

Multimodal LLM grounding molecular reasoning in chemical structures via Morgan fingerprint embeddings for drug discovery auditing.

Ax Sowmya Vajrala, Akshay Bankar, Manjunath Arveti, Shreyas Pandith, Sravanth Kodavanti, Subhajit Sanyal, Amit Unde, Srinivas Soumitri Miriyala 5/6/2026

TOC-SR: Task-Optimal Compact diffusion for Image Super Resolution

Compact one-step diffusion model for efficient image super-resolution via task-optimal backbone discovery.

Ax Ricardo Euler, Pedro Maristany de las Casas, Ralf Bornd\"orfer 5/6/2026

Logic-Constrained Shortest Paths for Flight Planning

Algorithm for logic-constrained shortest path problem with satisfiability constraints applied to flight planning optimization. Specialized constraint solving research.

Ax Hyunji Min, Sangwon Jung, Junyoung Sung, Dosung Lee, Leekyeung Han, Paul Hongsuck Seo 5/6/2026

GOAT: A Training Framework for Goal-Oriented Agent with Tools

GOAT training framework for LLM agents with tool use, synthesizes API execution data from documentation for fine-tuning without human annotation.

Ax Vojtech Franc, Jakub Paplham 5/6/2026

Epistemic Reject Option Prediction

Reject option prediction method quantifying both epistemic and aleatoric uncertainty for high-stakes prediction applications.

Ax Weihao Bo, Shan Zhang, Yanpeng Sun, Jingjing Wu, Qunyi Xie, Xiao Tan, Kunbin Chen, Wei He, Xiaofan Li, Na Zhao, Jingdong Wang, Zechao Li 5/6/2026

Agentic Learner with Grow-and-Refine Multimodal Semantic Memory

Agentic learner with multimodal semantic memory that grows and refines domain knowledge across multiple modalities to avoid repeating mistakes.

Ax Ali Ezzat Shahroor, Mohamed Bayan Kmainasi, Abul Hasnat, Dimitar Dimitrov, Giovanni Da San Martino, Preslav Nakov, Firoj Alam 5/6/2026

MemeLens: Multilingual Multitask VLMs for Memes

MemeLens is a multilingual multitask Vision Language Model for meme understanding across hate detection, propaganda, and sentiment tasks.

Ax Ruijie Shi, Houbin Zhang, Yuecheng Han, Yuheng Wang, Jingru Fan, Runde Yang, Yufan Dang, Huatao Li, Dewen Liu, Yuan Cheng, Chen Qian 5/6/2026

AgentXRay: White-Boxing Agentic Systems via Workflow Reconstruction

Introduces Agentic Workflow Reconstruction (AWR) task to synthesize explicit, interpretable workflows from opaque LLM-based agentic systems.

Ax Varun Ursekar (Emily), Apaar Shanker (Emily), Veronica Chatrath (Emily), Yuan (Emily), Xue, Sam Denton 5/6/2026

VeRO: An Evaluation Harness for Agents to Optimize Agents

Introduces VeRO, an evaluation harness for coding agents optimizing other agents through iterative edit-execute-evaluate cycles.