HN paulpauper 4/9/2026

When AI Day of Reckoning?

Analysis of AI ROI expectations and failure rates. Discusses coding as LLMs' most promising application area.

HN handfuloflight 4/9/2026

The Waymo Rule for AI-Generated Code

Article argues AI-generated code only needs to exceed developer capability, discusses formal methods and language design.

Ax Jianhong Pang, Ruoxi Cheng, Ziyi Ye, Xingjun Ma, Zuxuan Wu, Xuanjing Huang, Yu-Gang Jiang 4/9/2026

Steering the Verifiability of Multimodal AI Hallucinations

Framework for steering verifiability of multimodal LLM hallucinations, distinguishing between obvious and elusive hallucinations to guide mitigation strategies.

Ax Seongwoo Jeong, Seonil Son 4/9/2026

How Much LLM Does a Self-Revising Agent Actually Need?

Empirical study decomposing LLM-based agent competence to identify which capabilities derive from the language model versus explicit structural design in self-revising agents.

Ax Nguyen Phuc Tran, Brigitte Jaumard, Oscar Delgado, Tristan Glatard, Karthikeyan Premkumar, Kun Ni 4/9/2026

LLM-Augmented Knowledge Base Construction For Root Cause Analysis

Evaluates LLM-augmented knowledge base construction for root cause analysis in network communications to enable rapid failure diagnosis and outage resolution.

Ax Peijie Yu, Wei Liu, Yifan Yang, Jinjian Li, Zelong Zhang, Xiao Feng, Feng Zhang 4/9/2026

Benchmarking LLM Tool-Use in the Wild

Benchmark for evaluating LLM tool-use agents on multi-turn, multi-step interactions addressing compositional tasks, implicit intent, and instruction transitions in real user behavior.