Ax Shuzheng Si, Haozhe Zhao, Yu Lei, Qingyi Wang, Dingwei Chen, Zhitong Wang, Zhenhailong Wang, Kangyang Luo, Zheng Wang, Gang Chen, Fanchao Qi, Minjia Zhang, Maosong Sun 5/6/2026

From Context to Skills: Can Language Models Learn from Context Skillfully?

Arxiv paper on context learning in language models via inference-time skill augmentation for reasoning over complex contexts exceeding parametric knowledge.

Ax Zhensu Sun, Haotian Zhu, Bowen Xu, Xiaoning Du, Li Li, David Lo 5/6/2026

Towards Agentic Runtime Healing

Paper on using LLMs for automated runtime healing in self-healing systems, replacing predefined rules with adaptive error recovery.

Ax Xuan Shen, Yizhou Wang, Yufa Zhou, Xiangxi Shi, Pu Zhao, Yanzhi Wang, Jiuxiang Gu 5/6/2026

Efficient Reasoning with Hidden Thinking

Research on Heima framework that compresses chain-of-thought reasoning in MLLMs into abstract thinking tokens for efficiency.

Ax Alexandra Bazarova, Aleksandr Yugay, Andrey Shulga, Alina Ermilova, Andrei Volodichev, Konstantin Polev, Julia Belikova, Rauf Parchiev, Dmitry Simakov, Maxim Savchenko, Andrey Savchenko, Serguei Barannikov, Alexey Zaytsev 5/6/2026

Hallucination Detection in LLMs with Topological Divergence on Attention Graphs

TOHA detector for identifying LLM hallucinations in RAG systems by analyzing topological divergence patterns in attention graph structures.

Ax Jiongli Zhu, Yue Wang, Bailu Ding, Philip A. Bernstein, Vivek Narasayya, Surajit Chaudhuri 5/6/2026

MINT: Multi-Vector Search Index Tuning

MINT framework for tuning index selection strategies in multi-vector databases to optimize performance across multiple feature dimensions.

Ax Jiaqi Chen, Yanzhe Zhang, Yutong Zhang, Yijia Shao, Diyi Yang 5/6/2026

Generative Interfaces for Language Models

Proposes generative interfaces paradigm to move LLM interactions beyond linear request-response format for more efficient multi-turn, information-dense, and exploratory tasks.

Ax Lo\"ic Cabannes, Maximilian Beck, Gergely Szilvasy, Matthijs Douze, Maria Lomeli, Jade Copet, Pierre-Emmanuel Mazar\'e, Gabriel Synnaeve, Herv\'e J\'egou 5/6/2026

Short window attention enables long-term memorization

Analysis of hybrid attention architectures combining local sliding window and global attention for improved long-term memory.

Ax Zihan Wang, Zhongkui Ma, Xinguo Feng, Chuan Yan, Dongge Liu, Ruoxi Sun, Derui Wang, Minhui Xue, Guangdong Bai 5/6/2026

Re-Key-Free, Risky-Free: Adaptable Model Usage Control

Method for protecting deep neural networks from unauthorized use by enabling model usage control without embedding access keys in parameters.