Ax Tianyu Liu, Wangjie Zheng, Rui Yang, Benny Kai Guo Loo, Hui Zhang, Jeffries Lauran, Jianlei Gu, Botao Yu, Weihao Xuan, Kexin Huang, Nan Liu, James Zou, Yonghui Jiang, Hua Xu, Hongyu Zhao 5/8/2026

A Versatile AI Agent for Rare Disease Diagnosis and Risk Gene Prioritization

Hygieia: multi-modal AI agent for rare disease diagnosis integrating phenotypic, genetic, and clinical data sources.

Ax Xinquan Chen, Zhenyun Yin, Shan He, Bin Huang, Shanzhe Lei, Pengcheng Shi, Kun Cai, Bei Chen, Bangwei Liu, Zeyu Kang, Chao Huang, Yang Zhang, Wenjie Li, Ruijun Ge, Yajie Wang, Tianshun Fang, Tianyang Xu, Yiwen Cong, Meng Jin, Gaolei Li, Xuansheng Wu, Linhan Liu, Zijing He, An Li, Yan Teng, Xin Tan, ChaoChao Lu, Ji He, Jie Li, Chunfeng Song, Jinya Xu, Fan Song, Shujie Wang, Jianmin Qian, Jie Hou, Xuhong Wang, Yingchun Wang, Hui Wang, Xia Hu 5/8/2026

Safactory: A Scalable Agent Factory for Trustworthy Autonomous Intelligence

Safactory: infrastructure for scalable agent development covering evaluation, data management, and continuous improvement loops.

Ax Aleksandr Bowkis, Marie Davidsen Buhl, Jacob Pfau, Geoffrey Irving 5/8/2026

Automated alignment is harder than you think

Analysis of risks from using AI agents to automate alignment research, including potential for misleading safety assessments.

Ax Jamelle Watson-Daniels, Himaghna Bhattacharjee, Skyler Wang, Brandon Handoko, Antonio Li, Anaelia Ovalle, Mahesh Pasupuleti, Candace Ross, Vidya Sarma, Arjun Subramonian, Karen Ullrich, Will van der Vaart, Yijing Xin, Maximilian Nickel 5/8/2026

SCRuB: Social Concept Reasoning under Rubric-Based Evaluation

SCRuB framework for evaluating LLM reasoning about social concepts using rubric-based methodology.

Ax Siru Ouyang, Jun Yan, Yanfei Chen, Rujun Han, Zifeng Wang, Bhavana Dalvi Mishra, Rui Meng, Chun-Liang Li, Yizhu Jiao, Kaiwen Zha, Maohao Shen, Vishy Tirumalashetty, George Lee, Jiawei Han, Tomas Pfister, Chen-Yu Lee 5/8/2026

SkillOS: Learning Skill Curation for Self-Evolving Agents

Framework for LLM agents to learn and curate reusable skills from streaming tasks enabling self-evolution and continuous improvement.

Ax Tianle Wang, Zhaoyang Wang, Guangchen Lan, Xinpeng Wei, Sipeng Zhang, Guanwen Qiu, Abulhair Saparov 5/8/2026

Can RL Teach Long-Horizon Reasoning to LLMs? Expressiveness Is Key

Study of reinforcement learning for improving LLM long-horizon reasoning via controlled synthetic environment examining task difficulty and expressiveness.

Ax Daniel Zheng, Ingrid von Glehn, Yori Zwols, Iuliya Beloshapka, Lars Buesing, Daniel M. Roy, Martin Wattenberg, Bogdan Georgiev, Tatiana Schmidt, Andrew Cowie, Fernanda Viegas, Dimitri Kanevsky, Vineet Kahlon, Hartmut Maennel, Sophia Alj, George Holland, Alex Davies, Pushmeet Kohli 5/8/2026

AI Co-Mathematician: Accelerating Mathematicians with Agentic AI

Interactive workbench enabling mathematicians to leverage AI agents for exploratory research including literature search, computation, and theorem proving.

Ax Michael Timothy Bennett 5/8/2026

Are Flat Minima an Illusion?

Theoretical analysis questioning whether flat minima in loss landscapes causally explain generalization or are artifacts of parameterization.

Ax Tatiana Gaintseva, Andrew Stepanov, Ziquan Liu, Martin Benning, Gregory Slabaugh, Jiankang Deng, Ismail Elezi 5/8/2026

MidSteer: Optimal Affine Framework for Steering Generative Models

Theoretical framework for steering intermediate representations in generative models via affine transformations for alignment and safety.