Ax Gaurav Kamath, Sreenath Madathil, Sebastian Schuster, Marie-Catherine de Marneffe, Siva Reddy 3/2/2026

Humans and LLMs Diverge on Probabilistic Inferences

Study showing divergence between human and LLM behavior on probabilistic inference tasks requiring non-deterministic reasoning.

Ax Jielin Qiu, Jianguo Zhang, Zixiang Chen, Liangwei Yang, Ming Zhu, Juntao Tan, Haolin Chen, Wenting Zhao, Rithesh Murthy, Roshan Ram, Akshara Prabhakar, Shelby Heinecke, Caiming, Xiong, Silvio Savarese, Huan Wang 3/2/2026

AudioCapBench: Quick Evaluation on Audio Captioning across Sound, Music, and Speech

AudioCapBench: benchmark for evaluating audio captioning of multimodal LLMs across sound, music, speech with 1,000 samples and LLM-as-Judge evaluation.

Ax Hariz Yet, Nguyen Thanh Tam, Mao V. Ngo, Lim Yi Shen, Lin Wei, Jihong Park, Binbin Chen, Tony Q. S. Quek 3/2/2026

SLA-Aware Distributed LLM Inference Across Device-RAN-Cloud

System design for distributed LLM inference across device, RAN-edge, and cloud tiers with latency constraints for 5G embodied AI applications.

Ax Dongxu Zhang, Yiding Sun, Pengcheng Li, Yumou Liu, Hongqiang Lin, Haoran Xu, Xiaoxuan Mu, Liang Lin, Wenbiao Yan, Ning Yang, Chaowei Fang, Juanjuan Zhao, Jihua Zhu, Conghui He, Cheng Tan 3/2/2026

PointCoT: A Multi-modal Benchmark for Explicit 3D Geometric Reasoning

Multimodal benchmark evaluating MLLMs on explicit 3D geometric reasoning with point clouds, exposing geometric hallucinations.

Ax Oscar Hill, Mateo Espinosa Zarlenga, Mateja Jamnik 3/2/2026

Hierarchical Concept-based Interpretable Models

Hierarchical concept embedding models improving neural network interpretability through human-readable concept representations.