Ax Debeshee Das, Julien Piet, Darya Kaviani, Luca Beurer-Kellner, Florian Tram\`er, David Wagner 5/6/2026

Trojan Hippo: Weaponizing Agent Memory for Data Exfiltration

Trojan Hippo attack exploiting LLM agent memory systems for data exfiltration through dormant payloads planted via single untrusted tool interactions.

Ax Mario Koddenbrock, Christoph Lange, Robin Legner, Martin J\"ager, Martin K\"ogler, Mariano N. Cruz Bournazou, Peter Neubauer, Felix Biessmann, Erik Rodner 5/6/2026

RamanBench: A Large-Scale Benchmark for Machine Learning on Raman Spectroscopy

First large-scale reproducible benchmark for machine learning on Raman spectroscopy, standardizing evaluation across fragmented datasets and spectral models.

Ax Christopher Kelly, Angelica Chowdhury, Alexandra Campili, Bimpe Ayoola, Devin Barbour, Thomas Chen Dawson, Ze Shen Chin, Rokas Gipi\v{s}kis 5/6/2026

Principles and Guidelines for Randomized Controlled Trials in AI Evaluation

Framework for standardizing AI evaluation through randomized controlled trials, adopting validity principles from established experimental disciplines.

Ax Akash Bonagiri, Gerard Janno Anderias, Saee Patil, Angelina Lai, Devang Borkar, Gezheng Kang, Ishant Gandhi, Setareh Rafatirad, Houman Homayoun 5/6/2026

STABLEVAL: Disagreement-Aware and Stable Evaluation of AI Systems

STABLEVAL framework for disagreement-aware evaluation of AI systems that models annotator reliability and item ambiguity instead of using majority vote.

Ax Mingyu Luo, Zihan Zhang, Zesen Liu, Yuchong Xie, Zhixiang Zhang, Dung Hiu Hilton Yeung, Wai Ip Lai, Ping Chen, Ming Wen, Dongdong She 5/6/2026

When Alignment Isn't Enough: Response-Path Attacks on LLM Agents

Identifies security vulnerabilities in Bring-Your-Own-Key LLM agent architectures where malicious relays can tamper with aligned LLM responses post-generation.

Ax Karima Makhlouf, Lamiaa Basyoni, Syed Khaderi, Gabriel Marquez, Peter Sotomango, Mahmoud Awawdah, Sami Zhioua 5/6/2026

On the Privacy of LLMs: An Ablation Study

Unified ablation study of LLM privacy attacks (MIA, AIA, DEA, backdoor) across system factors, revealing behavior under common deployment conditions.

Ax \"Onder G\"urcan, Moharram Challenger 5/6/2026

LLM-enabled Social Agents

Framework for LLM-enabled social agents with grounding in roles, norms, intentions and contextual constraints for meaningful social interaction.

Ax Roberto Pietrantuono, Luca Giamattei, Stefano Russo, Julien Siebert, Neil Walkinshaw 5/6/2026

Causal Software Engineering: A Vision and Roadmap

Vision for causal inference methods in software engineering to support AI-driven decision-making and LLM-based agents with interventional and counterfactual reasoning.

Ax William Lehn-Schi{\o}ler, Magnus Ruud Kj{\ae}r, Phillip Hempel, Magnus Guldberg Pedersen, Rahul Thapa, Bryan He, Nicolai Spicher, Andreas Brink-Kjaer, Lars Kai Hansen, Emmanuel Mignot 5/6/2026

Pretraining on Sleep Data Improves non-Sleep Biosignal Tasks

Research investigating transfer learning from sleep biosignal pretraining to non-sleep medical tasks.