Ax Maksim Anisimov (Imperial College London), Francesco Belardinelli (Imperial College London), Matthew Wicker (Imperial College London) 4/13/2026

SafeAdapt: Provably Safe Policy Updates in Deep Reinforcement Learning

Method for safely updating deep reinforcement learning policies while preserving safety guarantees on previously encountered tasks.

Ax Wenjie Qu, Xuandong Zhao, Jiaheng Zhang, Dawn Song 4/13/2026

Self-Sovereign Agent

Investigation of self-sovereign AI agents that can economically sustain themselves without human involvement using LLMs and agent frameworks.

Ax Dengjia Zhang, Alexander Martin, William Jurayj, Kenton Murray, Benjamin Van Durme, Reno Kriz 4/13/2026

Unified Multimodal Uncertain Inference

Multimodal inference task with text, audio, video for producing calibrated probability estimates of hypotheses with fine-grained uncertainty.

Ax Jinghan Zhang, Fengran Mo, Tharindu Cyril Weerasooriya, Ruimin Dai, Xiaoyan Han, Yanjie Fu, Dakuo Wang, Kunpeng Liu 4/13/2026

StaRPO: Stability-Augmented Reinforcement Policy Optimization

RL framework for improving LLM reasoning by optimizing for logical consistency and structural integrity of reasoning processes, not just final answers.