Ax Thanh Dat Hoang, Thanh Trung Huynh, Matthias Weidlich, Thanh Tam Nguyen, Tong Chen, Hongzhi Yin, Quoc Viet Hung Nguyen 5/7/2026

FINER-SQL: Boosting Small Language Models for Text-to-SQL

FINER-SQL boosts small language models for text-to-SQL tasks, enabling efficient on-premise deployment with improved reasoning and instruction following.

Ax John Yang, Kilian Lieret, Jeffrey Ma, Parth Thakkar, Dmitrii Pedchenko, Sten Sootla, Emily McMilin, Pengcheng Yin, Rui Hou, Gabriel Synnaeve, Diyi Yang, Ofir Press 5/7/2026

ProgramBench: Can Language Models Rebuild Programs From Scratch?

Benchmark testing language model agents' ability to build complete software projects from scratch with minimal human oversight.

Ax Sebastian Wind, Tri-Thien Nguyen, Jeta Sopa, Mahshad Lotfinia, Sebastian Bickelhaup, Michael Uder, Harald K\"ostler, Gerhard Wellein, Sven Nebelung, Daniel Truhn, Andreas Maier, Soroosh Tayebi Arasteh 5/7/2026

Safety and accuracy follow different scaling laws in clinical large language models

SaFE-Scale framework showing clinical LLM safety and accuracy follow different scaling laws; safety doesn't necessarily improve with model size or context.

Ax Ivaxi Sheth, Jan Wehner, Sahar Abdelnabi, Ruta Binkyte, Mario Fritz 5/7/2026

Safety Must Precede the Deployment of Open-Ended AI

Position paper on safety requirements for open-ended AI agents with autonomous behavior generation and self-evolution capabilities before deployment.

Ax Rick Chen, Joseph Ternasky, Afriyie Samuel Kwesi, Ben Griffin, Aaron Ontoyin Yin, Zakari Salifu, Kelvin Amoaba, Xianling Mu, Fuat Alican, Yigit Ihlamur 5/7/2026

VCBench: Benchmarking LLMs in Venture Capital

VCBench: First benchmark for predicting founder success in venture capital using LLMs, with sparse signals and uncertain outcomes exceeding market index performance.

Ax Wenyue Hua, Tianyi Peng, Chi Wang, Jiaxin Pei, Ian Kaufman, Bryan Lim, Chandler Fang 5/7/2026

Quantifying Trust: Financial Risk Management for Trustworthy AI Agents

Framework for quantifying trust in autonomous AI agents through end-to-end operational outcomes rather than model-internal properties, relevant to deployed agents with financial risk.

Ax Saad Alqithami 5/7/2026

Soft Tournament Equilibrium

Soft Tournament Equilibrium: Framework for evaluating non-transitive LLM-based agents using set-valued cores instead of linear rankings for cyclic competitive domains.

Ax Tu Trinh, Mohamed Elfeki, Guangze Luo, Kelvin Luu, Nathan Hunt, Ernesto Hernandez, Nandan Marwaha, Yannis Yiming He, Charles Wang, Fernando Carabedo, Alessa Castillo, Bing Liu 5/7/2026

HiL-Bench (Human-in-Loop Benchmark): Do Agents Know When to Ask for Help?

HiL-Bench evaluates whether coding agents know when to ask for help versus act autonomously on incomplete specifications, addressing judgment gaps in frontier agents.