Ax Reva Schwartz, Carina Westling, Morgan Briggs, Marzieh Fadaee, Isar Nejadgholi, Matthew Holmes, Fariza Rashid, Maya Carlyle, Afaf Ta\"ik, Kyra Wilson, Peter Douglas, Theodora Skeadas, Gabriella Waters, Rumman Chowdhury, Thiago Lacerda 3/2/2026

CIRCLE: A Framework for Evaluating AI from a Real-World Lens

CIRCLE: six-stage framework for evaluating AI systems under real-world conditions and user variability beyond model-centric metrics.

Ax Borja Requena Pozo, Austin Letson, Krystian Nowakowski, Izan Beltran Ferreiro, Leopoldo Sarra 3/2/2026

A Minimal Agent for Automated Theorem Proving

Minimal agentic baseline for automated theorem proving that enables systematic comparison across AI-based prover architectures with iterative refinement and library search.

Ax Xuanming Cui, Hong-You Chen, Hao Yu, Hao Yuan, Zihao Wang, Shlok Kumar Mishra, Hanchao Yu, Yonghuan Yang, Jun Xiao, Ser-Nam Lim, Jianpeng Cheng, Qi Guo, Xiangjun Fan 3/2/2026

Reason to Contrast: A Cascaded Multimodal Retrieval Framework

TTE-v2 hybrid multimodal retrieval framework extending reasoning-driven bi-encoder architectures with improved performance.

Ax Gaurav Kamath, Sreenath Madathil, Sebastian Schuster, Marie-Catherine de Marneffe, Siva Reddy 3/2/2026

Humans and LLMs Diverge on Probabilistic Inferences

Study showing divergence between human and LLM behavior on probabilistic inference tasks requiring non-deterministic reasoning.

Ax Jielin Qiu, Jianguo Zhang, Zixiang Chen, Liangwei Yang, Ming Zhu, Juntao Tan, Haolin Chen, Wenting Zhao, Rithesh Murthy, Roshan Ram, Akshara Prabhakar, Shelby Heinecke, Caiming, Xiong, Silvio Savarese, Huan Wang 3/2/2026

AudioCapBench: Quick Evaluation on Audio Captioning across Sound, Music, and Speech

AudioCapBench: benchmark for evaluating audio captioning of multimodal LLMs across sound, music, speech with 1,000 samples and LLM-as-Judge evaluation.

Ax Hariz Yet, Nguyen Thanh Tam, Mao V. Ngo, Lim Yi Shen, Lin Wei, Jihong Park, Binbin Chen, Tony Q. S. Quek 3/2/2026

SLA-Aware Distributed LLM Inference Across Device-RAN-Cloud

System design for distributed LLM inference across device, RAN-edge, and cloud tiers with latency constraints for 5G embodied AI applications.