Ax Siru Ouyang, Jun Yan, I-Hung Hsu, Yanfei Chen, Ke Jiang, Zifeng Wang, Rujun Han, Long T. Le, Samira Daruki, Xiangru Tang, Vishy Tirumalashetty, George Lee, Mahsan Rofouei, Hangfei Lin, Jiawei Han, Chen-Yu Lee, Tomas Pfister 3/18/2026

ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory

ReasoningBank memory framework enables LLM agents to learn from interaction history and distill generalizable reasoning strategies for continuous tasks.

Ax Sumanth Varambally, Marshall Fisher, Jas Thakker, Yiwei Chen, Zhirui Xia, Yasaman Jafari, Ruijia Niu, Manas Jain, Veeramakali Vignesh Manivannan, Zachary Novack, Luyu Han, Srikar Eranky, Salva R\"uhling Cachay, Taylor Berg-Kirkpatrick, Duncan Watson-Parris, Yi-An Ma, Rose Yu 3/18/2026

Zephyrus: An Agentic Framework for Weather Science

Zephyrus agentic framework combines weather foundation models with LLM reasoning for interactive scientific workflows in meteorology.

Ax Dachuan Lin, Guobin Shen, Zihao Yang, Tianrong Liu, Dongcheng Zhao, Yi Zeng 3/18/2026

Efficient LLM Safety Evaluation through Multi-Agent Debate

Multi-agent debate framework using small language models for cost-efficient LLM safety evaluation, with HAJailBench benchmark for jailbreak testing.

Ax Sunghyun Wee, Suyoung Kim, Hyeonjin Kim, Kyomin Hwang, Nojun Kwak 3/18/2026

Alignment-Aware Quantization for LLM Safety

Alignment-Aware Quantization: PTQ method for efficient LLM deployment that preserves behavioral alignment and safety properties, not just minimizing reconstruction error.

Ax Nuoya Xiong, Yuhang Zhou, Hanqing Zeng, Zhaorun Chen, Furong Huang, Shuchao Bi, Lizhu Zhang, Zhuokai Zhao 3/18/2026

Token-Level LLM Collaboration via FusionRoute

FusionRoute: token-level collaboration method enabling multiple specialized LLMs to work together, combining domain expertise efficiency with generalization.

Ax Ruoran Li, Xinghua Zhang, Haiyang Yu, Shitong Duan, Xiang Li, Wenxin Xiang, Chonghua Liao, Xudong Guo, Yongbin Li, Jinli Suo 3/18/2026

MemPO: Self-Memory Policy Optimization for Long-Horizon Agents

MemPO: self-memory policy optimization approach enabling long-horizon agents to proactively manage memory content aligned with task objectives.

Ax Linus Folkerts, Will Payne, Simon Inman, Philippos Giavridis, Joe Skinner, Sam Deverett, James Aung, Ekin Zorer, Michael Schmatz, Mahmoud Ghanem, John Wilkinson, Alan Steer, Vy Hong, Jessica Wang 3/18/2026

Measuring AI Agents' Progress on Multi-Step Cyber Attack Scenarios

Evaluation of frontier AI models' autonomous cyber-attack capabilities on multi-step scenarios, tracking capability trends across 18 months of model releases.

Ax Vishnu Narayanan Anilkumar, Abhijith Sreesylesh Babu, Trieu Hai Vo, Mohankrishna Kolla, Alexander Cuneo 3/18/2026

Relationship-Aware Safety Unlearning for Multimodal LLMs

Framework for unlearning relational safety failures in multimodal LLMs where combinations of benign concepts become unsafe when linked by specific relations.

Ax Yulin Peng, Xinxin Zhu, Chenxing Wei, Nianbo Zeng, Leilei Wang, Ying Tiffany He, F. Richard Yu 3/18/2026

SAGE: Multi-Agent Self-Evolution for LLM Reasoning

SAGE framework: multi-agent reinforcement learning system for improving LLM reasoning without large human-labeled datasets, using self-play and closed-loop feedback.

Ax Sebastian Stober, Tim W. Dornis 3/18/2026

Generative AI Training and Copyright Law

Interdisciplinary study on copyright law implications of training generative AI via web scraping, covering fair use and TDM exceptions.

Ax Donato Crisostomi, Alessandro Zirilli, Antonio Andrea Gargiulo, Maria Sofia Bucarelli, Simone Scardapane, Fabrizio Silvestri, Iacopo Masi, Emanuele Rodol\`a 3/18/2026

MASS: MoErging through Adaptive Subspace Selection

MASS method merges multiple fine-tuned models via adaptive subspace selection, improving accuracy over existing merging approaches without retraining.

Ax Eleonora Cappuccio (Department of Computer Science, University of Pisa), Andrea Esposito (Department of Computer Science, University of Bari Aldo Moro), Francesco Greco (Department of Computer Science, University of Bari Aldo Moro), Giuseppe Desolda (Department of Computer Science, University of Bari Aldo Moro), Rosa Lanzilotti (Department of Computer Science, University of Bari Aldo Moro), Salvatore Rinzivillo (ISTI CNR) 3/18/2026

Explanation User Interfaces: A Systematic Literature Review

Systematic literature review of explanation user interfaces for interpretable AI systems.

Ax Zhe Ye, Zhengxu Yan, Jingxuan He, Timothe Kasriel, Kaiyu Yang, Dawn Song 3/18/2026

VERINA: Benchmarking Verifiable Code Generation

VERINA benchmark for evaluating LLM code generation with joint code, specification, and proof generation.

Ax Weihua Du, Hailei Gong, Zhan Ling, Kang Liu, Lingfeng Shen, Xuesong Yao, Yufei Xu, Dingyuan Shi, Yiming Yang, Jiecao Chen 3/18/2026

Generalizable End-to-End Tool-Use RL with Synthetic CodeGym

CodeGym: Reinforcement learning framework for training LLM agents to use tools generalizing across new tasks and workflows.