Ax Ahmed Heakl, Gustavo Bertolo Stahl, Sarim Hashmi, Seung Hun Eddie Han, Mukul Ranjan, Arina Kharlamova, Salman Khan, Abdulrahman Mahmoud 4/22/2026

CASS: Nvidia to AMD Transpilation with Data, Models, and Benchmark

Dataset and ML models for transpiling GPU code between CUDA and HIP architectures, with 60k verified code pairs.

Ax Shashank Sharma, Janina Hoffmann, Vinay Namboodiri 4/22/2026

MRS: Multi-Resolution Skills for HRL Agents

Hierarchical RL approach using multi-resolution skills to improve manager subgoal selection and performance on agile tasks.

Ax Zhen Zhu, Yiming Gong, Yao Xiao, Yaoyao Liu, Derek Hoiem 4/22/2026

How to Teach Large Multimodal Models New Skills

Research on sequential fine-tuning of large multimodal models showing skill recovery across different tasks and model families.

Ax Qiushi Han, David Simchi-Levi, Renfei Tan, Zishuo Zhao 4/22/2026

Multi-agent Adaptive Mechanism Design

DRAM framework combines mechanism design and online learning for sequential multi-agent truthful reporting.

Ax Basab Jha, Firoj Paudel, Ujjwal Puri, Ethan Henkel, Zhang Yuting, Mateusz Kowalczyk, Mei Huang, Choi Donghyuk, Wang Junhao 4/22/2026

SAGE-32B: Agentic Reasoning via Iterative Distillation

SAGE-32B is a 32B parameter model fine-tuned via iterative distillation for agentic reasoning, task decomposition, and tool usage.

Ax Zongyue Qin, Raghavv Goel, Mukul Gagrani, Risheek Garrepalli, Mingu Lee, Yizhou Sun 4/22/2026

ConFu: Contemplate the Future for Better Speculative Sampling

ConFu improves speculative decoding for LLM inference acceleration by enhancing draft model quality to propose better candidate tokens for verification.

Ax Hongyi Jin, Bohan Hou, Guanjie Wang, Ruihang Lai, Jinqi Chen, Zihao Ye, Yaxing Cai, Yixin Dong, Xinhao Cheng, Zhihao Zhang, Yilong Zhao, Yingyi Huang, Lijie Yang, Jinchen Jiang, Gabriele Oliaro, Jianan Ji, Xupeng Miao, Vinod Grover, Todd C. Mowry, Zhihao Jia, Tianqi Chen 4/22/2026

Event Tensor: A Unified Abstraction for Compiling Dynamic Megakernel

Event Tensor abstraction eliminates kernel launch overheads in LLM inference by fusing operators into persistent kernels handling dynamic shapes.