Ax Geigh Zollicoffer, Minh Vu, Hongli Zhan, Raymond Li, Manish Bhattarai 5/12/2026

Sanity Checks for Long-Form Hallucination Detection

Methodology for detecting hallucinations in LLM chain-of-thought reasoning by distinguishing actual reasoning evaluation from answer-surface correlates.

Ax Junnan Liu, Linhao Luo, Thuy-Trang Vu, Gholamreza Haffari 5/12/2026

AIPO: : Learning to Reason from Active Interaction

AIPO method for improving LLM reasoning through active interaction and reinforcement learning, extending beyond policy model capability boundaries.

Ax Xuanqiang Angelo Huang, Charlie Tharas, Samuele Marro, Van Q. Truong, Bernhard Sch\"olkopf, Emanuele La Malfa, Zhijing Jin 5/12/2026

Mechanism Design Is Not Enough: Prosocial Agents for Cooperative AI

Research proving mechanism design alone is insufficient for safe AI agent cooperation, proposing prosocial agent approaches for beneficial multi-agent interaction.

Ax Ramon Pires, Thales Sales Almeida, Celio Larcher Junior, Giovana Bon\'as, Hugo Abonizio, Marcos Piau, Roseval Malaquias Junior, Thiago Laitz, Rodrigo Nogueira 5/12/2026

Magis-Bench: Evaluating LLMs on Magistrate-Level Legal Tasks

Magis-Bench benchmark for evaluating LLMs on magistrate-level legal judgment tasks including weighing claims and rendering reasoned decisions.

Ax Giuseppe Bruno, Shi Chen, Zhengjiang Lin, Yury Polyanskiy, Philippe Rigollet 5/12/2026

Scaling Limits of Long-Context Transformers

Analyzes softmax attention scaling limits to understand when selectivity emerges versus uniform averaging in long-context transformers.

Ax Annan Yu, Dongwei Lyu, N. Benjamin Erichson 5/12/2026

Continuity Laws for Sequential Models

Studies continuity as inductive bias in sequential models, analyzing whether continuous-time formulations like state-space models behave continuously.

Ax Hassan Shapourian, Kasra Hejazi, Olabode M. Sule, Beren Millidge 5/12/2026

ZAYA1-VL-8B Technical Report

Technical report on ZAYA1-VL-8B, a compact mixture-of-experts vision-language model achieving competitive performance with smaller parameter count.

Ax Liam Davis, Leopold Haller, Alberto Alfarano, Mark Santolucito 5/12/2026

Lattice Deduction Transformers

Lattice Deduction Transformer uses recurrent attention with lattice projections for logical reasoning tasks, achieving perfect accuracy on constraint satisfaction with minimal parameters.

Ax Md Atik Ahamed, Mihir Parmar, Palash Goyal, Chun-Liang Li, Qiang Cheng, Tomas Pfister, Jinsung Yoon 5/12/2026

Reasoning-Aware Training for Time Series Forecasting

STRIDE combines time series foundation models with LLM reasoning to improve forecasting interpretability while handling continuous numerical values without excessive tokenization.

Ax Yinwei Dai, Zhuofu Chen, Lijie Yang, Ravi Netravali 5/12/2026

Geometry Guided Self-Consistency for Physical AI

Proposes KeyStone, an inference-time self-consistency method using geometry guidance for physical AI models that generate action trajectories via diffusion/flow matching.

Ax Aritra Mazumder, Shubhashis Roy Dipta, Nusrat Jahan Lia, Tanzila Khan, Kainat Raisa Hossain, Nehaa Shri, Shubhrangshu Debsarkar, Humayra Tasnim, Gour Gupal Talukder Shawon, Debjoty Mitra, Sumaiya Ahmed Rani, Al Jami Islam Anik, Al Nafeu Khan 5/12/2026

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators

Introduces AgentCollabBench, a diagnostic benchmark for measuring multi-hop process failures in multi-agent systems where individual agents appear correct but collaboration silently fails.

Ax Chengcheng Sun, Chenhao Li, Xiang Lin, Tianji Zheng, Fanrong Meng, Xiaobin Rui, Zhixiao Wang 5/12/2026

Attention-based graph neural networks: a survey

Survey of attention-based graph neural networks covering mechanisms for adaptive feature selection and noise filtering in graph representation learning.

Ax Ziyun Liu, Fengmiao Bian, Jian-Feng Cai 5/12/2026

AdaPreLoRA: Adafactor Preconditioned Low-Rank Adaptation

Proposes AdaPreLoRA, an optimization method for Low-Rank Adaptation using Adafactor preconditioning to handle rank-deficiency issues in LoRA factor-space preconditioners.