Ax Yanna Jiang, Delong Li, Haiyu Deng, Baihe Ma, Xu Wang, Qin Wang, Guangsheng Yu 2/25/2026

SoK: Agentic Skills -- Beyond Tool Use in LLM Agents

Systematization of knowledge on agentic skills—reusable procedural capabilities with explicit conditions, policies, and interfaces for reliable long-horizon workflows in LLM agents.

Ax Alagappan Ramanathan, Eunju Kang, Dongsu Han, Sangeetha Abdu Jyothi 2/25/2026

Airavat: An Agentic Framework for Internet Measurement

Airavat: agentic framework automating internet measurement workflows and verification against methodological standards using AI agents for tool orchestration.

Ax Mark Marron 2/25/2026

Toward an Agentic Infused Software Ecosystem

Vision paper proposing Agentic Infused Software Ecosystem (AISE) framework rethinking software development tools and practices for autonomous AI agents.

Ax Yang Zhang, Danyang Li, Yuxuan Li, Xin Zhang, Tianyu Xie, Mingming Cheng, Xiang Li 2/25/2026

CrystaL: Spontaneous Emergence of Visual Latents in MLLMs

CrystaL method enables spontaneous emergence of visual latent representations in MLLMs through improved supervision for latent chain-of-thought reasoning.

Ax Rui Zhao, Xihui Li, Yizheng Zhang, Yuzhen Liu, Zhong Zhang, Yufeng Zhang, Cheng Zhou, Zhengyou Zhang, Lei Han 2/25/2026

Cooperative-Competitive Team Play of Real-World Craft Robots

Multi-agent deep RL system for training cooperative-competitive robot teams in simulation with transfer to real-world robotic applications.

Ax Jorge Gallego-Feliciano, S. Aaron McClendon, Juan Morinelli, Stavros Zervoudakis, Antonios Saravanos 2/25/2026

Hidden Dynamics of Massive Activations in Transformer Training

Comprehensive analysis of massive activation patterns in transformer training using Pythia model family with public dataset release.

Ax Tianshi Zheng, Kelvin Kiu-Wai Tam, Newt Hue-Nam K. Nguyen, Baixuan Xu, Zhaowei Wang, Jiayang Cheng, Hong Ting Tsang, Weiqi Wang, Jiaxin Bai, Tianqing Fang, Yangqiu Song, Ginny Y. Wong, Simon See 2/25/2026

NewtonBench: Benchmarking Generalizable Scientific Law Discovery in LLM Agents

NewtonBench: Benchmark for evaluating LLM agents in scientific law discovery addressing memorization, scalability, and authentic scientific process.

Ax Arnab Sen Sharma, Giordano Rogers, Natalie Shapira, David Bau 2/25/2026

LLMs Process Lists With General Filter Heads

Mechanistic analysis revealing that LLMs use compact filter head representations to encode general filtering operations in list-processing tasks.

Ax Huanyao Zhang, Jiepeng Zhou, Bo Li, Bowen Zhou, Yanzhe Shan, Haishan Lu, Zhiyong Cao, Jiaoyang Chen, Yuqian Han, Zinan Sheng, Zhengwei Tao, Hao Liang, Jialong Wu, Yang Shi, Yuanpeng He, Jiaye Lin, Qintong Zhang, Guochen Yan, Runhao Zhao, Zhengpin Li, Xiaohan Yu, Lang Mei, Chong Chen, Wentao Zhang, Bin Cui 2/25/2026

BrowseComp-$V^3$: A Visual, Vertical, and Verifiable Benchmark for Multimodal Browsing Agents

BrowseComp benchmark for evaluating multimodal browsing agents with visual verification and deep search capabilities in open-world environments.