HN skydiver7373 5/13/2026

How LLMs Work

Technical explanation of language model mechanics covering how transformer-based systems process and generate text.

Ax Min Yang, Jinghua Piao, Xu Xia, Xiaochong Lan, Jiaju Chen, Yongshun Gong, Yong Li 5/13/2026

SkillMaster: Toward Autonomous Skill Mastery in LLM Agents

SkillMaster: Framework enabling LLM agents to autonomously develop, refine, and internalize skills through experience rather than external governance.

Ax Kun Xiang, Terry Jingchen Zhang, Zirong Liu, Bokai Zhou, Yueling Tang, Junjie Yu, Jiacong Lu, Shangrui Huang, Heng Li, Likui Zhang, Kunkun Liu, Changzheng Zhang, Yangle Fang, Boqiang Guo, Hui-Ling Zhen, Dandan Tu, Yinya Huang, Xiaodan Liang 5/13/2026

SeePhys Pro: Diagnosing Modality Transfer and Blind-Training Effects in Multimodal RLVR for Physics Reasoning

SeePhys Pro: Benchmark testing modality transfer in multimodal models for physics reasoning with progressively increasing visual information.

Ax Songlin Bai, Xintong Wang, Linlin Yu, Bin Chen, Zhiang Xu, Yuyang Sheng, Changtong Zan, Xiaofeng Zhu, Yizhe Zhang, Jiru Li, Mingze Guo, Ling Zou, Yalong Li, Chengfu Huo, Liang Ding 5/13/2026

IndustryBench: Probing the Industrial Knowledge Boundaries of LLMs

IndustryBench: 2,049-item benchmark for LLM performance on industrial procurement QA in Chinese, evaluating safety-critical correctness beyond standard metrics.

Ax Stepan Kulibaba, Kirill Labzin, Artem Dzhalilov, Roman Pakhomov, Oleg Svidchenko, Alexander Gasnikov, Aleksei Shpilman 5/13/2026

SDG-MoE: Signed Debate Graph Mixture-of-Experts

Proposes SDG-MoE, a sparse mixture-of-experts model with expert communication via signed debate graphs to improve token routing performance.

Ax Gal Engelberg, Leon Goldberg, Konstantin Koutsyi, Boris Plotnikov, Tiltan Gilat, Ben Benhemo 5/13/2026

AI Native Asset Intelligence

Proposes AI-native security assistant for enterprise environments that proactively prioritizes fragmented security signals by exposure and exploitability.

Ax Roxana Geambasu, Mariana Raykova, Pierre Tholoniat, Trishita Tiwari, Lillian Tsai, Wen Zhang 5/13/2026

Engineering Robustness into Personal Agents with the AI Workflow Store

Argues for engineering robustness into AI agents through software engineering practices like iterative design, testing, and staged deployment instead of on-the-fly synthesis.

HN aosmith 5/13/2026

Show HN: Gremlin

Gremlin is a browser-native TypeScript multi-agent coordinator with Svelte UI, supporting local LLM provider integration without server requirement.