Ax Mehul Bafna, Siddhant anand Jadhav, David Sweet 5/7/2026

Taking the GP Out of the Loop

Scaling Bayesian optimization to many observations by replacing Gaussian process surrogates with more scalable alternatives.

Ax Yi Ru Wang, Carter Ung, Christopher Tan, Grant Tannert, Jiafei Duan, Josephine Li, Anh Le, Rishabh Oswal, Markus Grotz, Wilbert Pumacay, Yuquan Deng, Ranjay Krishna, Dieter Fox, Siddhartha Srinivasa 5/7/2026

RoboEval: Where Robotic Manipulation Meets Structured and Scalable Evaluation

RoboEval: structured evaluation framework and benchmark for robotic manipulation with behavioral and outcome metrics beyond binary success.

Ax Friederike Groschupp, Daniele Lain, Aritra Dhar, Lara Magdalena Lazier, Srdjan \v{C}apkun 5/7/2026

Can LLMs Make (Personalized) Access Control Decisions?

Study evaluating LLM capability for personalized access control decisions in applications and agent-based systems to reduce user cognitive burden.

Ax Ignacio Heredia, \'Alvaro L\'opez Garc\'ia, Fernando Aguilar G\'omez, Diego Aguirre, Caterina Alarc\'on Mar\'in, Khadijeh Alibabaei, Lisana Berberi, Miguel Caballer, Amanda Calatrava, Pedro Castro, Alessandro Costantini, Mario David, Jaime D\'iez Stefan Dlugolinsky, Borja Esteban Sanchis, Giacinto Donvito, Leonhard Duda, Sa\'ul Fernandez, Andr\'es Heredia Canales, Valentin Kozlov, Sergio Langarita, Jo\~ao Machado, Germ\'an Molt\'o, Daniel San Mart\'in, Martin \v{S}eleng, Giang Nguyen, Marcin P{\l}\'ociennik, Marta Obreg\'on Ruiz, Susana Rebolledo Ruiz, Vicente Rodriguez, Judith S\'ainz-Pardo D\'iaz, Viet Tran 5/7/2026

AI4EOSC: a Federated Cloud Platform for Artificial Intelligence in Scientific Research

Open-source federated platform for operationalizing AI/ML lifecycle in scientific research with FAIR principles and MLOps tooling.

Ax Mingzhe Lu, Yiwen Wang, Yanbing Liu, Qi You, Chong Liu, Ruize Qin, Haoyu Dong, Wenyu Zhang, Jiarui Zhang, Yue Hu, Yunpeng Li 5/7/2026

LitVISTA: A Benchmark for Narrative Orchestration in Literary Text

Benchmark dataset (LitVISTA) for evaluating LLM narrative structure and story arcs in literary text generation versus human narratives.

Ax Daniel Zhu, Zihan Wang, Xuchan Bao, Jerry Wei 5/7/2026

Jailbroken Frontier Models Retain Their Capabilities

Analysis showing advanced jailbreaks on frontier LLMs scale with model capability and increasingly impose minimal performance degradation ('jailbreak tax').