Ax Zhe Ye, Zhengxu Yan, Jingxuan He, Timothe Kasriel, Kaiyu Yang, Dawn Song 3/18/2026

VERINA: Benchmarking Verifiable Code Generation

VERINA benchmark for evaluating LLM code generation with jointly generated specifications and proofs, addressing correctness verification challenges.

Ax Yidi Wang, Ziyue Qiao, Jiawei Gu, Xubin Zheng, Pengyang Wang, Xiaobing Pei, Xiao Luo 3/18/2026

Out-of-Distribution Graph Models Merging

Graph model merging technique for combining GNN models pre-trained on different domains with distribution discrepancy to create generalized models.

Ax Weihua Du, Hailei Gong, Zhan Ling, Kang Liu, Lingfeng Shen, Xuesong Yao, Yufei Xu, Dingyuan Shi, Yiming Yang, Jiecao Chen 3/18/2026

Generalizable End-to-End Tool-Use RL with Synthetic CodeGym

Tool-augmented LLM agents trained with synthetic code environments via RL to improve generalization on tool-use tasks, addressing brittleness with new tools and unseen workflows.

Ax Piotr Komorowski, Elena Golimblevskaia, Reduan Achtibat, Thomas Wiegand, Sebastian Lapuschkin, Wojciech Samek 3/18/2026

Attribution-Guided Decoding

Attribution-Guided Decoding uses interpretability to improve LLM instruction-following and factual accuracy.

Ax Bahrul Ilmi Nasution, Floor Eijkelboom, Mark Elliot, Richard Allmendinger, Christian A. Naesseth 3/18/2026

Flow Matching for Tabular Data Synthesis

Empirical comparison of flow matching variants with diffusion models for privacy-preserving tabular data synthesis.

Ax Yunni Qu (The University of North Carolina at Chapel Hill), Dzung Dinh (The University of North Carolina at Chapel Hill), Grant King (University of Michigan), Whitney Ringwald (University of Minnisota Twin Cities), Bing Cai Kok (The University of North Carolina at Chapel Hill), Kathleen Gates (The University of North Carolina at Chapel Hill), Aidan Wright (University of Michigan), Junier Oliva (The University of North Carolina at Chapel Hill) 3/18/2026

Relaxed Efficient Acquisition of Context and Temporal Features

Active feature acquisition method for biomedical applications optimizing measurement selection under temporal and cost constraints.

Ax Vincent Zhihao Zheng, \'Etienne Marcotte, Arjun Ashok, Andrew Robert Williams, Lijun Sun, Alexandre Drouin, Valentina Zantedeschi 3/18/2026

Overcoming the Modality Gap in Context-Aided Forecasting

Research addressing multimodal model underperformance in context-aided forecasting via improved context quality assessment.

Ax Seth Karten, Jake Grigsby, Tersoo Upaa Jr, Junik Bae, Seonghun Hong, Hyunyoung Jeong, Jaeyoon Jung, Kun Kerdthaisong, Gyungbo Kim, Hyeokgi Kim, Yujin Kim, Eunju Kwon, Dongyu Liu, Patrick Mariglia, Sangyeon Park, Benedikt Schink, Xianwei Shi, Anthony Sistilli, Joseph Twin, Arian Urdu, Matin Urdu, Qiao Wang, Ling Wu, Wenli Zhang, Kunsheng Zhou, Stephanie Milani, Kiran Vodrahalli, Amy Zhang, Fei Fang, Yuke Zhu, Chi Jin 3/18/2026

The PokeAgent Challenge: Competitive and Long-Context Learning at Scale

Large-scale benchmark for AI agents combining partial observability, game-theoretic reasoning, and long-horizon planning in Pokemon battle environment.

Ax Xiao Zhu, Chenmien Tan, Pinzhen Chen, Rico Sennrich, Huiming Wang, Yanlin Zhang, Hanxu Hu 3/18/2026

CHARM: Calibrating Reward Models With Chatbot Arena Scores

CHARM method calibrating reward models using Chatbot Arena scores to mitigate model preference bias, improving alignment of LLMs through RLHF.

Ax Mathew J. Koretsky, Maya Willey, Owen Bianchi, Chelsea X. Alvarado, Tanay Nayak, Nicole Kuznetsov, Sungwon Kim, Mike A. Nalls, Daniel Khashabi, Faraz Faghri 3/18/2026

BiomedSQL: Text-to-SQL for Scientific Reasoning on Biomedical Knowledge Bases

BiomedSQL benchmark for text-to-SQL generation requiring scientific reasoning over biomedical knowledge bases, evaluating LLM capability for complex analytical tasks.