Ax Eleonora Cappuccio (Department of Computer Science, University of Pisa), Andrea Esposito (Department of Computer Science, University of Bari Aldo Moro), Francesco Greco (Department of Computer Science, University of Bari Aldo Moro), Giuseppe Desolda (Department of Computer Science, University of Bari Aldo Moro), Rosa Lanzilotti (Department of Computer Science, University of Bari Aldo Moro), Salvatore Rinzivillo (ISTI CNR) 2/20/2026

Explanation User Interfaces: A Systematic Literature Review

Systematic literature review of explanation user interfaces for black-box AI systems and XAI techniques.

Ax Yan Wang, Lingfei Qian, Xueqing Peng, Yang Ren, Keyi Wang, Yi Han, Dongji Feng, Fengran Mo, Shengyuan Lin, Qinchuan Zhang, Kaiwen He, Chenri Luo, Jianxing Chen, Junwei Wu, Chen Xu, Ziyang Xu, Jimin Huang, Guojun Xiong, Xiao-Yang Liu, Qianqian Xie, Jian-Yun Nie 2/20/2026

FinTagging: Benchmarking LLMs for Extracting and Structuring Financial Information

FinTagging benchmark for evaluating LLMs on financial information extraction and hierarchical GAAP concept classification.

Ax Thibaud Gloaguen, Robin Staab, Nikola Jovanovi\'c, Martin Vechev 2/20/2026

Watermarking Diffusion Language Models

First watermarking scheme designed for diffusion language models that generate tokens in arbitrary order rather than sequentially.

Ax Luca Belli, Kate Bentley, Will Alexander, Emily Ward, Matt Hawrilenko, Kelly Johnston, Mill Brown, Adam Chekroud 2/20/2026

VERA-MH Concept Paper

VERA-MH automated evaluation framework for assessing safety of AI chatbots in mental health contexts using LLM-based agents.

Ax Ahmed Aboulfotouh, Hatem Abou-Zeid 2/20/2026

Multimodal Wireless Foundation Models

Wireless foundation models extended to process multiple modalities for improved task performance across varying operating conditions.

Ax Mozes Jacobs, Thomas Fel, Richard Hakim, Alessandra Brondetta, Demba Ba, T. Andy Keller 2/20/2026

Block-Recurrent Dynamics in Vision Transformers

Block-Recurrent Hypothesis characterizes Vision Transformer depth as block-recurrent structure, providing mechanistic understanding of ViT computations.

Ax Marie S. Bauer, Julia Gachot, Matthias Kerzel, Cornelius Weber, Stefan Wermter 2/20/2026

Theory of Mind for Explainable Human-Robot Interaction

Framework integrating Theory of Mind into robots for inferring human mental states to enhance explainability and predictability in human-robot interaction.

Ax Cen Zhang, Younggi Park, Fabian Fleischer, Yu-Fu Fu, Jiho Kim, Dongkwan Kim, Youngjoon Kim, Qingxiao Xu, Andrew Chin, Ze Sheng, Hanqing Zhao, Brian J. Lee, Joshua Wang, Michael Pelican, David J. Musliner, Jeff Huang, Jon Silliman, Mikel Mcdaniel, Jefferson Casavant, Isaac Goldthwaite, Nicholas Vidovich, Matthew Lehman, Taesoo Kim 2/20/2026

SoK: DARPA's AI Cyber Challenge (AIxCC): Competition Design, Architectures, and Lessons Learned

Analysis of DARPA's AIxCC competition for autonomous cyber reasoning systems leveraging LLMs to discover vulnerabilities in open-source software.

Ax Zhiliang Chen, Alfred Wei Lun Leong, Shao Yong Ong, Apivich Hemachandra, Gregory Kang Ruey Lau, Chuan-Sheng Foo, Zhengyuan Liu, Nancy F. Chen, Bryan Kian Hsiang Low 2/20/2026

The Chicken and Egg Dilemma: Co-optimizing Data and Model Configurations for LLMs

Framework for jointly optimizing data mixture and model architecture configurations during LLM training through co-optimization rather than sequential approaches.

Ax Iv\'an Arcuschin, David Chanin, Adri\`a Garriga-Alonso, Oana-Maria Camburu 2/20/2026

Biases in the Blind Spot: Detecting What LLMs Fail to Mention

Automated black-box pipeline detects unverbalized biases in LLM reasoning where models hide internal biases in plausible-sounding chain-of-thought explanations.

Ax Beatrix M. G. Nielsen, Emanuele Marconato, Luigi Gresele, Andrea Dittadi, Simon Buchholz 2/20/2026

Logit Distance Bounds Representational Similarity

Proves logit distance bounds representational similarity for discriminative models including autoregressive language models.

Ax Md. Najib Hasan, Touseef Hasan, Souvika Sarkar 2/20/2026

Are LLMs Ready to Replace Bangla Annotators?

Evaluates LLMs as zero-shot annotators for Bangla hate speech detection, examining reliability and bias in low-resource language settings.

Ax Nils Palumbo, Sarthak Choudhary, Jihye Choi, Prasad Chalasani, Somesh Jha 2/20/2026

Policy Compiler for Secure Agentic Systems

PCAS system enforces deterministic authorization policies in LLM agents for customer service, workflows, and compliance without relying on prompts.

Ax Xidong Wang, Shuqi Guo, Yue Shen, Junying Chen, Jian Wang, Jinjie Gu, Ping Zhang, Lei Liu, Benyou Wang 2/20/2026

LiveClin: A Live Clinical Benchmark without Leakage

LiveClin live benchmark for clinical LLM evaluation using contemporary peer-reviewed cases updated biannually to prevent contamination.