Ax Yunni Qu (The University of North Carolina at Chapel Hill), Dzung Dinh (The University of North Carolina at Chapel Hill), Grant King (University of Michigan), Whitney Ringwald (University of Minnisota Twin Cities), Bing Cai Kok (The University of North Carolina at Chapel Hill), Kathleen Gates (The University of North Carolina at Chapel Hill), Aidan Wright (University of Michigan), Junier Oliva (The University of North Carolina at Chapel Hill) 3/18/2026

Relaxed Efficient Acquisition of Context and Temporal Features

Active feature acquisition method for biomedical applications optimizing measurement selection under temporal and cost constraints.

Ax Vincent Zhihao Zheng, \'Etienne Marcotte, Arjun Ashok, Andrew Robert Williams, Lijun Sun, Alexandre Drouin, Valentina Zantedeschi 3/18/2026

Overcoming the Modality Gap in Context-Aided Forecasting

Research addressing multimodal model underperformance in context-aided forecasting via improved context quality assessment.

Ax Seth Karten, Jake Grigsby, Tersoo Upaa Jr, Junik Bae, Seonghun Hong, Hyunyoung Jeong, Jaeyoon Jung, Kun Kerdthaisong, Gyungbo Kim, Hyeokgi Kim, Yujin Kim, Eunju Kwon, Dongyu Liu, Patrick Mariglia, Sangyeon Park, Benedikt Schink, Xianwei Shi, Anthony Sistilli, Joseph Twin, Arian Urdu, Matin Urdu, Qiao Wang, Ling Wu, Wenli Zhang, Kunsheng Zhou, Stephanie Milani, Kiran Vodrahalli, Amy Zhang, Fei Fang, Yuke Zhu, Chi Jin 3/18/2026

The PokeAgent Challenge: Competitive and Long-Context Learning at Scale

Large-scale benchmark for AI agents combining partial observability, game-theoretic reasoning, and long-horizon planning in Pokemon battle environment.

Ax Xiao Zhu, Chenmien Tan, Pinzhen Chen, Rico Sennrich, Huiming Wang, Yanlin Zhang, Hanxu Hu 3/18/2026

CHARM: Calibrating Reward Models With Chatbot Arena Scores

CHARM method calibrating reward models using Chatbot Arena scores to mitigate model preference bias, improving alignment of LLMs through RLHF.

Ax Mathew J. Koretsky, Maya Willey, Owen Bianchi, Chelsea X. Alvarado, Tanay Nayak, Nicole Kuznetsov, Sungwon Kim, Mike A. Nalls, Daniel Khashabi, Faraz Faghri 3/18/2026

BiomedSQL: Text-to-SQL for Scientific Reasoning on Biomedical Knowledge Bases

BiomedSQL benchmark for text-to-SQL generation requiring scientific reasoning over biomedical knowledge bases, evaluating LLM capability for complex analytical tasks.

Ax Sumanth Varambally, Thomas Voice, Yanchao Sun, Zhifeng Chen, Rose Yu, Ke Ye 3/18/2026

Hilbert: Recursively Building Formal Proofs with Informal Reasoning

Hilbert system recursively building formal proofs by combining informal LLM reasoning with Lean 4 verification, bridging gap between mathematical reasoning and formal proof generation.

Ax Sumanth Varambally, Marshall Fisher, Jas Thakker, Yiwei Chen, Zhirui Xia, Yasaman Jafari, Ruijia Niu, Manas Jain, Veeramakali Vignesh Manivannan, Zachary Novack, Luyu Han, Srikar Eranky, Salva R\"uhling Cachay, Taylor Berg-Kirkpatrick, Duncan Watson-Parris, Yi-An Ma, Rose Yu 3/18/2026

Zephyrus: An Agentic Framework for Weather Science

Zephyrus framework combining foundation models for weather forecasting with LLM reasoning to enable language-based scientific workflows on meteorological datasets.

Ax L\'eo Boisvert, Abhay Puri, Chandra Kiran Reddy Evuru, Nazanin Sepahvand, Nicolas Chapados, Quentin Cappart, Alexandre Lacoste, Krishnamurthy Dj Dvijotham, Alexandre Drouin 3/18/2026

Malice in Agentland: Down the Rabbit Hole of Backdoors in the AI Supply Chain

Security research on backdoor attacks in AI agent supply chains through poisoned interaction data collection, formalizing threat models for finetuned web browsing and tool-use agents.

Ax Dashti A. Ali, Aras T. Asaad, Jacob J. Peoples, Ahmad Bashir Barekzai, Camila Vilela, Hala Khasawneh, Jayasree Chakraborty, Jo\~ao Miranda, Mohammad Hamghalam, Natalie Gangai, Natally Horvat, Richard K. G. Do, Alice C. Wei, Amber L. Simpson 3/18/2026

A Novel Patch-Based TDA Approach for Computed Tomography Imaging

Topological data analysis patch-based approach for CT imaging feature extraction improving ML model performance on medical diagnosis tasks.

Ax Danxu Liu, Di Wang, Hebaixu Wang, Haoyang Chen, Wentao Jiang, Yilin Cheng, Haonan Guo, Wei Cui, Jing Zhang 3/18/2026

SARMAE: Masked Autoencoder for SAR Representation Learning

Noise-aware masked autoencoder for self-supervised SAR satellite imagery representation learning addressing data scarcity and speckle noise challenges.

Ax Nuoya Xiong, Yuhang Zhou, Hanqing Zeng, Zhaorun Chen, Furong Huang, Shuchao Bi, Lizhu Zhang, Zhuokai Zhao 3/18/2026

Token-Level LLM Collaboration via FusionRoute

FusionRoute enables token-level collaboration between specialized and general-purpose LLMs via dynamic routing, improving efficiency and domain performance.

Ax Kai Wittenmayer, Sukrut Rao, Amin Parchami-Araghi, Bernt Schiele, Jonas Fischer 3/18/2026

CFM: Language-aligned Concept Foundation Model for Vision

Language-aligned concept foundation model decomposing vision representations into human-interpretable concepts with spatial grounding across diverse tasks.

Ax Linus Folkerts, Will Payne, Simon Inman, Philippos Giavridis, Joe Skinner, Sam Deverett, James Aung, Ekin Zorer, Michael Schmatz, Mahmoud Ghanem, John Wilkinson, Alan Steer, Vy Hong, Jessica Wang 3/18/2026

Measuring AI Agents' Progress on Multi-Step Cyber Attack Scenarios

Evaluates autonomous cyber-attack capabilities of frontier AI models on multi-step attack scenarios, comparing seven models over 18 months at varying inference compute budgets.

HN mimbojimbo 3/18/2026

GSD 2

GSD 2 is a standalone CLI coding agent built on Pi SDK, evolving from a Claude prompt framework to a full agent with session and context control.