Ax Mohit Talreja, Joshua Diao, Jim Thannikary James, Radu Casapu, Tejas Santanam, Ethan Mendes, Alan Ritter, Wei Xu, James Hays 4/21/2026

GeoRC: A Benchmark for Geolocation Reasoning Chains

Benchmark for evaluating vision-language model reasoning: assesses whether VLMs can explain geolocation predictions with supporting image evidence.

Ax Chenxi Wang, Zhuoyun Yu, Xin Xie, Wuguannan Yao, Runnan Fang, Shuofei Qiao, Kexin Cao, Guozhou Zheng, Xiang Qi, Peng Zhang, Shumin Deng 4/21/2026

SkillX: Automatically Constructing Skill Knowledge Bases for Agents

SkillX automatically constructs plug-and-play skill knowledge bases for LLM agents to improve learning efficiency and generalization.

Ax Yuxiang Wang, Hongyu Liu, Yijiang Xu, Qinke Ni, Li Wang, Wan Lin, Kunyu Feng, Dekun Chen, Xu Tan, Lei Wang, Jie Shi, Zhizheng Wu 4/21/2026

VoxSafeBench: Not Just What Is Said, but Who, How, and Where

VoxSafeBench introduces a benchmark for evaluating safety in speech language models within multi-user environments, considering speaker identity and environmental context.

Ax Jiamei Wu, Ce Zhang, Zhipeng Cai, Jingsen Kong, Bei Jiang, Linglong Kong, Lingchen Kong 4/21/2026

Differentially Private Conformal Prediction

Differentially private conformal prediction method for uncertainty quantification with statistical efficiency under privacy constraints.

BL 4/21/2026

Scaling Codex to enterprises worldwide

OpenAI expands Codex to enterprises through partnerships with GSIs, reaching 4M weekly developers. Codex deployment across software development lifecycle workflows.

HN _doctor_love 4/20/2026

AI and the Neurodivergent Unlock

Study examining how AI tools like Cursor enable improved workflow and focus for neurodivergent developers.

HN ChicagoDave 4/20/2026

DevArch 2.0

DevArch 2.0 provides directives and agents for Claude Code enforcing engineering discipline with automated quality gates and testing.